Planning, tools, and context: what actually changes coding-agent performance?
Learn how planning, tools, and context management affect coding agents and why the right harness depends on the model, task, and resource budget.
The model does not work alone. Planning, action interfaces, and context management shape the agent trajectory, but none of these components is an automatic improvement.
When a coding agent solves a difficult task, it is tempting to attribute the result to the model. When it fails, the first reaction is often to switch models or increase the context window.
That explanation is incomplete. Between the request and the code sits a layer that defines which instructions the agent receives, which tools it can use, how it tracks a plan, which observations remain in context, how edits are validated, and when execution should stop.
That layer is the coding harness: the operational structure that turns a model’s general capabilities into a software-engineering process. Two agents using the same model can behave differently because they expose different interfaces, context policies, validations, and continuity rules.
What the study tried to separate
The preprint An Empirical Study of Harness Design for Coding Agents caught my attention because, instead of comparing complete products, it tries to isolate three decisions: persistent planning, action space, and context management across long-running tasks.
The authors built a lightweight harness with a fixed execution loop. Permission handling, post-edit diagnostics, and stuck detection remained constant while the three components were changed separately.
- Four models: Nemotron-3 30B, 120B, and 550B, plus Mistral-Medium-3.5-128B.
- Two long-horizon benchmarks: SWE-Bench Verified and Terminal-Bench 2.1.
- Five context-management strategies.
- 32k, 64k, 96k, and 128k context-window budgets.
- 176 experimental settings with measurements of success, cost, and trajectory behavior.
As of October 7, 2026, the work remained a v1 preprint submitted to arXiv on September 17. It does not cover the entire agent landscape; it provides controlled evidence about specific implementations. That distinction prevents a conditional result from becoming a universal rule.
Context management helps the agent keep working
A long task accumulates instructions, file reads, search results, logs, errors, diffs, and test output. Even when each step is useful, the history can grow until it consumes the available window.
- `T0`: no additional compaction; execution stops when history exceeds the window.
- `T1`: elision of stale tool observations.
- `T2`: elision with external storage and on-demand recovery.
- `T3`: LLM summarization without prior elision.
- `T4`: staged elision, recoverable storage, and summarization.
Elision means replacing an old, bulky output with a short marker. Under T4, summarization is only used when this deterministic reduction is not enough.
The strongest benefit appeared under tighter windows. Much of the gain came from preventing context overflow from ending the run early. That does not necessarily mean the agent reasoned better; in many cases, it simply managed to reach editing and verification.
In the study, eliding stale observations before invoking summarization provided the strongest aggregate efficiency among the evaluated strategies.
Recoverable memory only helps when the agent uses it
Making removed observations recoverable sounds like an obvious improvement. However, the study found very little use of the recovery mechanism and no accuracy gain over elision alone in the tested settings.
This does not prove that external memory is useless. It shows that exposing a tool does not guarantee that the agent will recognize when it needs it. The agent must notice that an observation matters again, locate the right item, and use retrieval to change its next decision.
Planning can scaffold execution or reduce waste
Planning in the study was not a generic request for the model to think before acting. It was a persistent representation of progress, updated through a tool and reinjected at every turn.
For the weakest evaluated model, planning helped sustain the trajectory until an edit attempt, improving success at additional cost. For stronger models, its main effect was reducing repeated post-edit verification, with small accuracy changes.
- Current hypothesis.
- Relevant files and evidence found.
- Change applied.
- Pending validation.
- Next step.
A useful plan records working state. If it does not change decisions, reduce repetition, or help resume work, it may be nothing more than additional context to carry.
Predefined tools versus Bash has no universal winner
Typed tools expose operations such as reading, searching, and editing files through structured arguments. A Bash-based interface can combine several operations in one command and may reduce the number of interactions.
In the study, predefined tools helped models with weaker Bash proficiency. Models that were more capable in that environment worked effectively through Bash and, especially on terminal-centric tasks, combined operations at lower cost.
This does not mean strong models do not need tools. The comparison changed the complete interface: predefined tools also provided instructions, state tracking, write validation, and automatic diagnostics. The experiment did not isolate tool count or action granularity alone.
- Capability: can the model express the intended action correctly?
- Safety: do permissions and boundaries remain explicit?
- Observability: is it possible to understand what ran and why?
- Efficiency: how many interactions and tokens were required?
Bash can be a powerful interface without becoming unrestricted authorization. Typed tools can offer valuable protection without fragmenting every operation into calls that are too small.
Example: investigating a failure with extensive logs
Imagine an agent fixing an integration failure in a large repository. It must locate the flow, reproduce the error, edit the code, and run a test suite that emits thousands of lines.
1locate files2 -> reproduce the failure3 -> record a hypothesis4 -> edit code5 -> run tests6 -> inspect the remaining failure7 -> adjust the implementation8 -> verify againThree different problems may look like one model failure: the agent abandons the investigation before editing; old test results consume the window; or the interface forces so many small calls that cost grows before verification.
- Outcome: was the task completed and validated?
- Termination: did the agent stop because of completion, a limit, context, an error, or repetition?
- Progress: did it reach localization, reproduction, editing, and verification?
- Context: which outputs consumed the most space?
- Plan: did persistent state change decisions or prevent repetition?
- Actions and cost: how many calls, failures, and tokens were required?
I would not change everything at once. If most trajectories die because of context pressure, I would test eliding old logs before summarization. If the agent stops before the first edit, I would evaluate a persistent plan. If there are too many calls and the model is shell-proficient, I would compare a more composable interface without removing safety controls.
A simple framework for evaluating the harness
1. Define the task and baseline
Select representative tasks, success criteria, and an initial configuration. Without a baseline, every change can look like an improvement.
2. Instrument the trajectory
Record stages, tools, context size, termination reason, validations, retries, and cost. The final answer alone does not explain where the system failed.
3. Classify the bottleneck
- Abandonment before useful action.
- Tool or permission failure.
- Context overflow.
- Incorrect edit.
- Missing or repetitive verification.
- Declared completion without enough evidence.
4. Change one observable component
Test planning, context policy, or action interface in isolation whenever possible. If they all change together, the result becomes another comparison between black boxes.
5. Compare success and cost per completed task
An intervention may improve success while increasing cost. Another may reduce calls without preserving quality. The decision must focus on accepted outcomes, not only token count or number of steps.
What I would take from the study
- More context does not eliminate the need for context management; it only postpones the boundary.
- Planning, memory, and tools should address identified bottlenecks instead of becoming rituals.
- Model capability and interface design interact; useful support for one model may add friction for another.
- Trajectory analysis reveals more than final success alone.
Limits and caveats
- The work is a preprint, not a definitive consensus.
- It evaluates four models, including three sizes from the same family.
- The benchmarks are SWE-Bench Verified and Terminal-Bench 2.1; the former is limited to Python.
- Planning and action space were tested only under T4 and a 128k window.
- Each setting was run once per task, and Terminal-Bench contains 89 tasks.
- Model size is an imperfect proxy for capability.
- The tool comparison bundles schemas, prompts, state tracking, and diagnostics.
- Cost depends on the models, pricing, and experimental conditions.
The reported crossover points should not be transferred directly to another model, repository, or harness. They are useful hypotheses that still need validation in the target environment.
Conclusion
Planning, tools, and context management shape how an agent turns capability into completed work. A good harness is not the one with the most layers, but the one that makes the bottleneck observable, applies an appropriate intervention, and lets us verify whether success, cost, and safety actually improved.
Before switching models or adding more memory, I would start with the trajectory: where the agent stops, what fills the context, which actions fail, and which steps repeat. That analysis usually offers a more precise answer about which part of the system needs to change.
