OMP: the tools around the model

Inside Oh My Pi: snapshot-based edits, syntax-tree rewrites, persistent eval runtimes, typed subagents, and rules that interrupt generation.

In this note

A coding agent can understand a bug and still struggle to change the right line. It can produce a sensible refactor and miss a caller. It can explain a failing test without ever running it. Somewhere between the model's answer and the working code, there is a lot of software doing the less glamorous work.

Oh My Pi, usually shortened to OMP and launched as omp, makes that software the interesting part. It's a terminal coding agent with a broad set of development tools built in. To understand its appeal, follow a small change from a request, through an edit, to the evidence that the change works.

OMP's terminal interface: the welcome screen, model and workspace status, and prompt composer. Local capture of OMP 18.6.1 on 4 October 2026, cropped to the interface. The example prompt hasn't been submitted. Select the image for a full-size view.

What the harness does

OMP is an open-source fork of Pi, Mario Zechner's coding-agent toolkit. Its repository describes a terminal agent extended with native tooling, code intelligence, debugging, and subagents. The project is MIT-licensed.

Suppose a request says: “Allow three retries after the first request, then check the boundary cases.” The model needs to locate the retry code, understand how attempts are counted, change it, and inspect the result. OMP's job is to make those steps available within one conversation. Its overview describes this continuous repository-to-verification workflow.

This distinction helps explain why two agents using the same model can behave differently. One may give it a large file dump and an awkward patch format. Another may expose the relevant function, a precise edit operation, and a diagnostic with the location of a broken caller. Those are different working conditions for the same reasoning task.

Choose the model, keep the workflow

OMP separates its interface from the model provider. Its provider guide covers API keys, supported account sign-ins, local engines such as Ollama, and custom endpoints. Inside a session, /model opens the model picker. Availability depends on the credentials and endpoints configured for that installation.

There are also model roles. The main conversation can use one model, lightweight work another, and planning another. Roles such as default, smol, slow, and plan describe where a model is used. A specialist agent adds its own instructions and tool access on top of that choice.

The practical attraction is continuity. A team can keep its repository instructions and way of reviewing changes while trying a different model. That doesn't make models interchangeable in quality. Tool use, reasoning, latency, and cost still need to be judged on the work they actually perform.

One particularly interesting routing option is prewalk. Start with omp --model @slow --prewalk-into @smol, and the initial model can explore the code and establish the work before handing the same session to the target model. The first file edit or write triggers the one-time handoff. Reading alone doesn't. History and todo state stay with the session; this isn't a separate worker producing a plan that another agent has to reconstruct.

That makes a specific cost trade-off possible: spend stronger reasoning on discovery, then try a lighter model for execution. It also gives the handoff a sharp limitation. The first edit isn't necessarily the end of the hard reasoning. If the difficult part is interpreting a failing integration test after the change, switching models at that point may be premature. Prewalk is a routing mechanism to evaluate, not a promise that cheaper execution preserves the initial model's judgment.

Hashline gives an edit a target the tool can check

Editing is a surprisingly fussy interface problem. With string replacement, the model often has to reproduce the old text exactly before supplying the new text. A whitespace mismatch can turn a correct intended change into a failed operation. Line numbers alone have another problem: they can move after the file changes.

OMP's default edit mode is called Hashline, with model-specific fallbacks and other modes also available. In the documented read format, a mutable file has numbered lines and a four-character hexadecimal snapshot tag derived from its normalized content. The tag identifies the version the model saw.

Here's our retry example. Assume attempts start at 1. The current condition permits another request after attempts 1 and 2, giving three attempts in total. Three retries after the first request require four attempts.

A precise edit / illustrative tool exchange

Same file. Same snapshot. One changed line.

Read result[retry.ts#A1B2]
1:export function canRetry(attempt: number) {
2:  return attempt < 3;
3:}
Edit request[retry.ts#A1B2]
PUT 2.=2:
+  return attempt < 4;

Replace line 2 in the tagged snapshot. Supply only the new content.

A1B2

If the file changes, OMP must validate or safely recover the target. An unresolved mismatch returns an error.

The tags are illustrative, not computed from this snippet. The patch uses the current snapshot-tag syntax; the agent copies the real tag from its tool output.

PUT 2.=2: replaces the inclusive range from line 2 to line 2. The + row carries the replacement text. The model-facing edit instructions require the observed tag and original line positions, so several changes can address one snapshot without manually adjusting later line numbers.

The edit reference adds an important detail: stale tags can trigger recovery using recorded snapshots. If a unique safe recovery can't be established, the operation fails. The tool also checks which lines were actually exposed to the model, rather than allowing an edit into an omitted part of a structural summary by default.

That's a useful mechanical check, but it can't establish that 4 is the right value. The number depends on the counting convention and the requirement. The boundary cases still matter: canRetry(3) should allow the fourth attempt; canRetry(4) should stop.

The read format can also be inspected directly with omp read retry.ts, without asking a model to act. This local capture uses the exact three-line fixture above. Its computed tag is 40C7; A1B2 in the diagram is a placeholder for explaining the patch syntax.

OMP returns the snapshot header retry.ts#40C7 and three numbered lines, including return attempt less than 3 on line 2.
Actual local omp read retry.ts output from OMP 18.6.1, captured on 4 October 2026 and cropped to the terminal result. The read reference documents this snapshot format. The capture verifies the read result; no edit was executed for this image. Scroll horizontally on a narrow screen to inspect the full line.

A language server knows which name you mean

A precise text edit is enough for the comparison above. Renaming the function across a project raises a different question: which occurrences of canRetry refer to this function?

OMP exposes the Language Server Protocol, or LSP. A language server supplies the definitions, references, types, and diagnostics that power many editor features. OMP can ask it for a semantic rename, instead of asking the model to approximate one with a text search. The code-intelligence guide explains discovery and configuration.

The distinction becomes visible when two functions share a name. In this illustrative module layout, only references to the exported retry helper belong to the rename:

One spelling / two symbols

Matching text doesn't always mean matching code.

retry.tsexport function canRetry(…)Rename definition
client.tsimport { canRetry } from './retry';Rename reference
database.tsfunction canRetry(…)Different local function
Selected symbol becomeshasRetryBudget
The declaration and its import change together. The unrelated function keeps its name. This is a schematic example of symbol identity, not an OMP session capture.

This depends on an installed, working server and a correctly discovered project. It also has limits: dynamic string lookups and consumers outside the server's workspace may be invisible. The compiler and relevant tests remain useful after the rename.

Automatic feedback has separate settings. The LSP tool reference describes diagnostics and server operations. The documented defaults enable diagnostics after whole-file writes, while diagnostics after incremental edits and format-on-write are off. A tool existing in the harness doesn't mean every possible check runs after every edit.

Rewrite code shapes with ast_edit

Text positions and symbol identity cover two kinds of change. A repetitive migration needs another: find every call with a particular syntactic shape and transform it. OMP exposes ast_grep and ast_edit for this. An abstract syntax tree, or AST, represents parsed code as nodes such as calls, arguments, and declarations. Matching those nodes avoids treating formatting as the meaning of the code.

Imagine replacing calls to scheduleRetry with a new budget-aware API. This illustrative ast_edit request preserves the original arguments:

{
  "ops": [{
    "pat": "scheduleRetry($$$ARGS)",
    "out": "scheduleRetryWithBudget($$$ARGS)"
  }],
  "paths": ["packages/client/src/**/*.ts"]
}

The $$$ARGS metavariable captures zero or more syntax nodes. A single-dollar capture such as $ARG matches one node. The pattern reference explains these bindings; the editing guide shows how search and rewrite fit together. Calls split over several lines can still match the call shape. Comments containing the same spelling aren't call nodes.

This is syntax-aware rather than symbol-aware. A same-named function from a different module can still match. Scope the paths, inspect imports, and use LSP when binding identity matters. The AST answers “does this code have this structure?”; the language server answers “which declaration does this reference resolve to?” Those questions are related, but neither substitutes for the other.

ast_edit first returns a preview with applied: false, match counts, affected files, and parse problems. The agent accepts it by writing a reason to xd://resolve, or discards it through xd://reject. That is the tool's proposal lifecycle; it doesn't imply a person has approved the change. The implementation reference also documents an important edge case: resolution reruns the rewrite on current files, and checks for changed match counts after applying it. A stale-preview error can therefore accompany files that have already changed. Inspect the resulting diff instead of assuming an error means nothing was written.

Files with parse errors can be skipped, so a clean-looking match count doesn't prove every file was examined successfully. Structural search is also disabled by default through astGrep.enabled, while structural editing is enabled by default through astEdit.enabled. The distinction matters when reproducing a workflow from a list of advertised tools.

When the answer is in a running process

Some bugs survive a careful source read. Perhaps the retry counter resets between requests, or a wrapper calls the function with zero-based attempts. The comparison looks reasonable, but the value arriving at it is wrong.

OMP's debugger integration uses the Debug Adapter Protocol, or DAP. With the appropriate adapter installed, it can launch or attach to a program, set breakpoints, inspect variables and stack frames, and step through execution. The adapter provides the connection to the particular runtime.

For the retry example, a useful request would be: “Stop in the retry decision on the third failed request and inspect the attempt counter and its caller.” This asks for evidence at the decision point. It gives the investigation somewhere specific to look when the test result and the apparent source logic disagree.

The setup is real: the adapter must be available, and a compiled target may need debug symbols. A paused program also stops doing normal work until it resumes. Debugging adds access to runtime evidence, with the responsibilities that come from controlling a process.

Keep computation in a persistent runtime

Once the investigation involves many fixtures or a large log, sending every record through the conversation becomes wasteful. OMP's eval tool gives the model a persistent Python or JavaScript execution environment. It can retain a dataset, compute over it, and return only the rows or summary needed for the next decision. The distinctive part is that this runtime can also call the agent's tools.

The eval reference specifies one cell per invocation, selected by language: "py" or language: "js". Python uses an IPython-style kernel; JavaScript uses a retained Bun worker. The runtimes are separate. A JavaScript variable doesn't automatically become a Python variable, and disabling one backend doesn't silently redirect its code to the other.

For a hypothetical JSON fixture with attempt and allowed fields, the JavaScript cells could look like this:

// Cell 1: keep the full fixture in the runtime.
const cases = JSON.parse(await read("fixtures/retry-cases.json"));

// Cell 2, in a later eval call: show only disagreements.
const mismatches = cases.filter(row =>
  row.allowed !== (row.attempt < 4)
);
display({ checked: cases.length, mismatches });

read and display are eval helpers in this example. The data remains in the worker between cells. The model sees a bounded result rather than having to count rows in a huge transcript. That shifts deterministic filtering into code and leaves the model to interpret the discrepancy. These are illustrative cells, not a claim that a fixture or benchmark was run for this article.

Call tools from code

The tool bridge makes the runtime more than a scratchpad. For example, await tool.read({ path: "src/retry.ts" }) invokes the actual read tool. JavaScript can coordinate independent requests with Promise.all; Python can await tool calls on the kernel's persistent event loop. The calls still pass through the live tool registry and its approval policy. Writing a loop doesn't bypass an action's permissions.

OMP also documents model helpers, child-agent handles, and kernel-defined tools inside eval. This makes small orchestration programs possible: discover candidates, dispatch bounded investigations, collect results, then display a compact comparison. A child has its own eval executor, but a granted tool defined in the caller's kernel executes back in that caller's kernel. Workspace isolation and execution-state ownership are distinct.

For an MCP tool, structured output has a specific location: result.details.structuredContent. Check result.hasError before using it. The text preview can be truncated while structured content remains available, but OMP doesn't manufacture structured content when a server omits it or validate it against that server's output schema. Pagination and shape checks still belong in the calling code. This distinction prevents a short display from being mistaken for the complete data contract.

Persistent isn't restartable

Live runtime state and saved conversation state have different lifetimes. Compaction can include a bounded snapshot describing a live kernel's environment and loaded paths; it doesn't serialize every variable value. Resuming a session in a new process doesn't resurrect the old kernel. Likewise, editing a previously loaded file doesn't rerun it, and JavaScript imports can remain cached. Reload deliberately or reset the selected runtime when the investigation requires fresh state.

This is a useful engineering constraint rather than a small footnote. A result derived from cases is only current if that dataset is current. For repeatable validation, keep the durable computation in a script or test and use eval to explore it. The interactive kernel is excellent at carrying an investigation forward; a checked-in program is easier to reproduce after the session disappears.

Follow the process and the browser

The execution layer has more depth than passing every command to a fresh system shell. OMP's native shell architecture uses Rust bindings around a brush shell runtime and supports persistent shell instances. Many utilities, including search, text processing, and file operations, run as in-process builtins. External programs still execute as external programs. This puts process control, cancellation, and common command behavior in a layer the harness owns, without making every command identical to its system-binary counterpart.

Long-running work has an explicit lifecycle. A finite command can run in the foreground or become a managed background job. A named service can declare readiness conditions and expose status and logs through proc://<name>. The bash reference documents stdin, stop, and lifetime controls on these handles. Readiness and completion are different events: a web server can be ready to accept requests while its process continues running. That is exactly the state needed for a browser-based check.

OMP's browser API is an eval prelude rather than a standalone agent tool. With JavaScript eval and browser support enabled, code can open a tab, inspect it, operate elements, and collect page errors. This hypothetical local application exposes its retry count with a test attribute:

const tab = await browser.open({
  name: "retry-ui",
  url: "http://127.0.0.1:3000",
  viewport: { width: 390, height: 844 }
});
display(await tab.text("[data-retry-count]"));
display(await tab.errors());

An observation can also produce element IDs or accessibility references for subsequent actions. Those references belong to the observed page state; re-observe after a rerender rather than assuming an old handle still identifies the same control. Advanced tab.run calls execute in a separate per-tab worker, so variables in the eval kernel aren't automatically available inside the callback. Pass inputs explicitly. These details are ordinary automation concerns, and exposing them makes the interface more useful than a generic “browser tool” label.

For the retry change, source inspection establishes the comparison, a test establishes the counting contract, DAP can expose the actual counter, and the browser can show the state the user sees. Each observation closes a different gap. Successfully editing TypeScript doesn't establish that the retry button recovers correctly after the last permitted request.

Parallel work needs clear ownership

A one-line boundary fix has little to gain from several agents. A change to the retry policy across an API, a client library, and documentation may have independent pieces. OMP provides child agent sessions for that split.

The subagent guide makes the defaults worth checking: workers share the parent's checkout unless isolation is enabled. Isolation uses a platform-selected clone or overlay, requires Git, and can integrate successful changes as patches. It reduces simultaneous checkout interference; integration can still conflict.

An example split after the policy is agreed
Owner Responsibility Evidence returned
Implementation worker Retry helper and its callers Diff and boundary checks
Documentation worker Explain the agreed attempt count Updated example and wording
Main session Integrate the results Combined behaviour and final diff

Agreeing on “three retries means four attempts” comes first. Otherwise the workers can produce two locally sensible interpretations of an unsettled requirement. Parallel execution cannot resolve that dependency by itself.

Agent Hub, opened with Alt+A, exposes worker activity and usage, with controls to inspect and steer sessions. Each worker has its own model conversation, so additional workers can add requests and cost. Separate workspaces also aren't a security sandbox: tool access can still reach external services.

Make worker results a contract

A worker's answer needn't be an unstructured paragraph. The task tool accepts an outputSchema for each task, using JSON Schema to describe the required result. For a read-only boundary investigation, an illustrative batch request could require file locations and an explicit answer about the third attempt:

{
  "context": "Attempts begin at 1; three retries means four attempts.",
  "tasks": [{
    "name": "RetryBoundary",
    "task": "Inspect the retry helper and tests. Return evidence; do not edit.",
    "solutionSpace": "Open investigation: determine behavior from the helper and tests",
    "outputSchema": {
      "type": "object",
      "required": ["files", "thirdAttemptAllowsRetry"],
      "properties": {
        "files": { "type": "array", "items": { "type": "string" } },
        "thirdAttemptAllowsRetry": { "type": "boolean" }
      },
      "additionalProperties": false
    },
    "schemaMode": "strict"
  }]
}

The task contract distinguishes strict and permissive validation. Permissive is the default and can accept an invalid payload with a warning after schema retries are exhausted. Strict mode fails instead. Select it when downstream code depends on the shape. It verifies that thirdAttemptAllowsRetry is a boolean, not that the worker's conclusion about the implementation is true. Evidence still needs to support the field's value.

solutionSpace describes how open the problem is: whether a known change should be carried out or its cause still needs investigating. For a child using automatic reasoning-effort selection, that field is the classifier's input. It describes uncertainty, not just the number of files. Giving a narrow task an accurate scope can therefore affect both the work assigned and the reasoning effort chosen.

Full results have agent:// handles rather than depending on the inline summary fitting into context. JSON-path reads such as agent://<id>/files/0 can extract one field from structured output. The assigned ID comes from the actual tool result. This turns delegation into a contract a program can consume: bounded task, explicit result shape, retrievable artifact, and a validation status that must be checked.

Catch a bad edit while it's being generated

Standing instructions consume context even when they aren't relevant. A more unusual OMP feature is TTSR, Time Traveling Stream Rules. A rule waits for a pattern in the assistant's output or tool arguments, then injects its guidance when the pattern appears. In the default interrupting mode, that can stop the response before the matching tool call executes and retry with the missing instruction in context.

Suppose the retry policy must remain typed, and an agent starts silencing a type error with as any. Put this illustrative rule in .omp/rules/typed-retries.md:

---
description: Keep retry policy checks typed
condition: '\bas\s+any\b'
scope:
  - tool:edit(*.ts)
  - tool:write(*.ts)
---
Preserve the retry policy's types. Narrow unknown input or model the
missing field explicitly before changing the boundary condition.

The rule guide defines the condition as a JavaScript regular expression. Regex matching uses accumulated scoped output, so a pattern can match even when its text arrives across several streaming chunks. For edits, OMP extracts introduced source content instead of matching raw patch JSON. File-specific scopes also keep a TypeScript rule from inspecting an unrelated Markdown hunk in the same operation.

TTSR / an illustrative intercepted edit

The reminder arrives at the offending expression.

  1. Generating
    const policy = input as any;

    The introduced TypeScript matches the rule. Interrupt this response.

  2. Injecting
    Keep retry policy checks typed

    Narrow unknown input or model the missing field explicitly.

  3. Retrying
    const policy = parsePolicy(input);

    A possible corrected expression. Its parser and behavior still need checking.

The first matching edit hasn't executed in this example. A rule changes the next generation's context; it doesn't supply a verified implementation of parsePolicy.

The sequence is what makes the feature interesting. OMP detects a particular mistake, discards the interrupted partial message by default, adds the rule body as corrective context, and schedules a guarded retry. It records which rule fired so the same rule doesn't continuously retrigger. The injection lifecycle documents the default once-per-session behavior and the configurable repeat policy. Already completed actions are not rolled back.

There is also an AST condition option, but its timing is different. Regex rules can watch streamed content; syntax-tree rules inspect finalized, validated tool arguments before execution. They aren't incrementally parsing each incoming fragment. Passive rules, configured with interruptMode: never, let the action run and add their reminder afterward. That is a useful review mechanism, but it cannot prevent the action it has just observed.

A pattern rule is best treated as a targeted correction mechanism. The default fires once, a regex can miss an equivalent expression, and a syntactic condition doesn't establish intent. Use approval policies or execution hooks for a policy that must be enforced every time. For the rule above, OMP's test command can check the matcher without waiting for a model to happen to produce the mistake:

omp ttsr test --rule .omp/rules/typed-retries.md \
  --source tool --tool edit --path src/retry.ts \
  'const policy = input as any;'
OMP 18.6.1 tests the TypeScript edit snippet const policy equals input as any and reports one triggered rule, typed-retries.
Actual local matcher test using the rule body above, with OMP 18.6.1 on 4 October 2026. The rule file was named typed-retries.md in an isolated fixture directory. Cropped to the terminal output. The green Triggered (1) result verifies the condition and scope for this input; this test doesn't exercise a live model interruption or retry. Scroll horizontally on a phone for the full output.

Give the session an independent reviewer

TTSR responds to recognizable patterns. Some mistakes require a judgment: the agent has tested attempts 1 and 2 while claiming to have verified the boundary at 3 and 4. OMP's optional advisor attaches a reviewer model with its own context and tools to inspect the main agent's work. It receives transcript updates and can investigate the repository before sending advice back.

The advisor guide separates three severities: a nit is a non-interrupting aside, a concern can steer the work when delivery conditions allow, and a blocker can interrupt or prompt a follow-up even after an ordinary final answer. Deliberately stopping the main agent suppresses automatic restarts. Plan-mode review has its own delivery constraints. This is an asynchronous feedback channel, not an unconditional second model taking command.

A project WATCHDOG.md can give reviewer-specific guidance, for example: “Check whether the reported boundary cases were actually executed. Flag a claim about three retries if the initial attempt wasn't included in the count.” This guidance belongs to the reviewer context rather than being another ordinary instruction for the main agent. WATCHDOG.yml can define a roster of reviewers with different models, cadence, and tool grants.

The advisor architecture gives each reviewer a distinct tool session, including separate file snapshots and seen-line tracking. Default investigative access includes read, grep, and glob; optional mutating tools remain subject to the normal approval policies. The reviewer neither approves the primary agent's actions nor grants it new permissions. Its notes are advice the primary must weigh.

Use /advisor on to enable review for a session and /advisor status to inspect model, context, usage, and cost. A persistent configuration instead uses advisor.enabled and the advisor model role. More review adds model calls and can arrive late. It is most valuable when given a specific failure to look for, with the underlying tests and diff still available for inspection.

Teach the project, keep the session

Repository knowledge is easy to lose between tasks. OMP loads project context files such as AGENTS.md, which can state build commands, generated-file boundaries, and conventions. That gives each new request a practical starting point.

For a retry library, a useful instruction might be: “Attempts are numbered from 1. Public documentation counts retries after the first request.” That small fact prevents a recurring ambiguity more effectively than repeating a general request to be careful.

Sessions save completed conversation entries and tool activity automatically. Return to the project and run omp --continue to continue recent work, or omp --resume to choose a saved session. A resumed conversation retains the investigation's context; its assumptions about the files still need to be checked against the current checkout.

These mechanisms serve different purposes. Project instructions carry standing conventions. Session history carries the path taken through a particular problem. Keeping both useful means deciding which facts belong in the repository and which belong only to that investigation.

Long conversations also need a bounded working context. OMP's compaction machinery records a summary entry and a boundary indicating which recent entries to keep. Rebuilding model input combines the latest summary, the retained portion, and later entries. The persisted session and the messages sent on a particular model request therefore aren't identical. A saved transcript can preserve details the model no longer has verbatim in its active context.

That distinction connects back to eval and delegation. Durable facts should have durable handles: a test file, a saved output artifact, a recorded convention. A summary saying “the fixture was checked” is less useful than a reproducible command and the location of its result. Context maintenance helps a long investigation continue; it shouldn't become the only storage system for evidence.

A small first session

On macOS, the repository lists Homebrew as one installation option. The quickstart covers the first-run provider and model setup:

brew install can1357/tap/omp
omp --version
cd path/to/your-project
omp --approval-mode always-ask

The final flag makes writes and executable actions stop for confirmation while ordinary reads can proceed. The approval reference documents yolo as the built-in default, so choose the policy deliberately. An approval policy governs whether actions run; it doesn't make an allowed command harmless.

After connecting a provider, begin with a question whose answer can be checked:

Find where this project decides whether to retry a request.
Explain how attempts are counted and identify the boundary tests.
Do not edit yet. Cite the files and functions you used.

Then request the bounded change, the relevant check, and the diff. For this example, inspect the expected behaviour at attempts 3 and 4 before looking at how confidently the agent describes its work. That sequence makes a first session useful even if the initial hypothesis turns out to be wrong.

Embed the harness without the terminal UI

OMP's terminal interface is only one host for its session machinery. The project documents an SDK for embedding sessions, an Agent Client Protocol bridge for editor clients, and a stdio RPC mode. These interfaces make the harness useful as a component of another application while retaining tool execution and session events.

Start omp --mode rpc and communicate over newline-delimited JSON. This is OMP's own JSONL protocol, not JSON-RPC 2.0. A host can submit a request like this illustrative prompt frame:

{"id":"retry-check","type":"prompt","message":"Inspect retry counting. Do not edit."}

The RPC reference makes acceptance, completion, and quiescence separate states. A successful command response can mean the prompt was admitted. Its later prompt_result reports the prompt's outcome. session_settled means pending asynchronous work is no longer able to wake the conversation. A background worker can finish after the main agent yields, so treating the first acknowledgment as the final answer would be a protocol bug.

An embedding needs to keep reading events, correlate request IDs, handle tool and UI requests, and decide which session state its own interface displays. Interactive defaults also don't all carry over: RPC and ACP use host-oriented defaults for settings such as advisors and memory unless explicitly configured. The appeal is that these contracts are exposed. Building a custom interface doesn't require scraping terminal text and guessing whether the agent has finished.

Judge the change the tools help you see

What makes OMP compelling is the amount of development machinery it puts within reach of one conversation. A file can be an observed snapshot, a call can be a syntax node, a name can be a resolved symbol, and a runtime value can be inspected at a breakpoint. Data can remain in an eval worker while only the useful result enters context. A child can return a typed artifact. A streamed mistake can trigger a correction at the moment it appears. Those are concrete improvements to the information and operations available to the model.

Their guarantees are deliberately different. Hashline can check the edit's target; an AST can check its syntactic shape; a schema can check a worker result's structure. None establishes that “three retries” was interpreted correctly. That last claim belongs to the requirement and behavioral evidence. In our example, the proof is still straightforward: attempts begin at 1, failure at attempt 3 permits attempt 4, and failure at attempt 4 stops. If the interface also presents a retry control, its final state needs checking in the running application.

More machinery also means more choices to manage: tool availability, model routing, kernel lifetime, worker ownership, review cadence, and what happens after an error. The useful response is to make those choices serve the task. A tiny boundary fix may need one precise edit and two checks. A library-wide migration can justify structural rewrites, isolated workers, and an independent reviewer. The feature count matters less than whether the chosen mechanisms remove a real source of uncertainty.

The opening question was how a plausible model answer becomes working code. OMP gives that transition a rich, inspectable implementation. Start with a behavior whose contract can be stated, trace how the agent gathers evidence and changes it, then reproduce the check outside the conversation. A strong result leaves something the developer can review and run again: a clear diff, a tested boundary, and enough context to understand why the change is right.

Keep this article in your starred list.

↑ ↓ to explore · Enter to open