Akshath Tiwari

This is a deep dive on SkillOpt, from Yang and colleagues at Microsoft Research, arXiv 2605.23904, a technique that treats prompt optimization as gradient descent over text. The central idea is to give text optimization the discipline of weight optimization. In ordinary gradient descent you have parameters theta, a loss, a gradient that points toward higher loss, and a learning rate that controls step size. You cannot differentiate a block of text, so SkillOpt makes an analogy: the parameter is a single skill document, a persistent natural-language instruction file that is the frozen agent’s external state; the loss is one minus the task score r of s, where r is the harness score of running the agent with that skill; the gradient is a separate optimizer language model that reads full execution traces, splits them into successes and failures, and emits bounded add, delete, and replace edits on the document, a textual gradient in the lineage of ProTeGi from EMNLP 2023 and TextGrad; and the learning rate is an edit budget L sub t, the maximum number of edits applied per step, which is ranked and clipped and follows a schedule, by default a cosine decay from four edits down to a floor of two. The discipline that makes this reliable, rather than a loosely self-revising loop that regresses as often as it helps, is a strict validation gate: every candidate skill is scored on a held-out selection split and accepted only when it is strictly greater than the current skill’s score, so ties are rejected and the skill never drifts sideways on noise; rejected edits drop into a buffer that acts as momentum, a negative memory that stops the optimizer re-proposing dead ends. On top of this fast per-step update path runs a slow update path once per epoch that compares the skill at the start versus the end of the epoch, categorizes improvements, regressions, persistent failures, and stable successes, and writes longitudinal guidance into a protected region of the document that step edits cannot overwrite, plus an optimizer-side meta-skill memory that never ships. Because the optimizer is a training-time tool and what deploys is just a short text file, SkillOpt adds zero inference-time model calls at deployment, and the learned skills stay compact at 379 to 1995 tokens after only one to four accepted edits. Results, all from the paper: across six benchmarks, seven target models, and three execution harnesses of direct chat, Codex, and Claude Code, SkillOpt is best or tied best on all 52 evaluated cells and beats every per-cell competitor including human-written skills, one-shot LLM skills, Trace2Skill, TextGrad, GEPA, and EvoSkill. On GPT-5.5 in direct chat it lifts the six-benchmark average by 23.5 points, by 24.8 inside Codex, and by 19.1 inside Claude Code; per benchmark on GPT-5.5 direct chat the no-skill to SkillOpt scores were SearchQA 77.7 to 87.3, SpreadsheetBench 41.8 to 80.7, OfficeQA 33.1 to 72.1, DocVQA 78.8 to 91.2, LiveMath 37.6 to 66.9, and ALFWorld 83.6 to 95.5. The honest caveats are that the gradient is a metaphor and not calculus since nothing is differentiated and the learning rate is a heuristic edit count, that gains are largest where the base model is weakest so it raises floors more than ceilings, that the zero inference cost applies only at deployment while training spends many rollouts and a strong optimizer model, that it needs a scorable task, and that the results come from the authors’ own suite and count ties. Two interactive widgets let the reader compare a too-small, an overshooting, and a cosine edit-budget schedule to feel why the middle path wins, and step through a single optimization step of run, grade, propose a textual-gradient edit, and validation gate, watching accepted edits stick and a rejected step roll back while the held-out score ratchets upward. This post is a companion to the GEPA deep dive, one of SkillOpt’s baselines.