vr
← Writing

Mop First, Prompt Later: Claude Context Custodian

A practical case for treating context as a budget, and keeping the useful signal close.

In this note

The Server Doesn’t Remember You

Every time you send a message to Claude, the request hits a stateless API server. That server has no memory of you. It doesn’t know you asked it something five minutes ago. It doesn’t know your name, your codebase, or what you were working on yesterday.

What it does know is exactly what you send it — right now, in this request.

This is the foundational fact of working with AI models: the context window is the only memory the model has. There is no hidden state on the server side. No user profile being looked up. No prior conversation being retrieved. Just the tokens in the current request, processed once, response returned, state discarded.

This is why Claude Code doesn’t just send your message to the model — it sends your message plus the entire conversation history, plus the system prompt, plus every file it read, plus every tool result. All of it, every turn. The continuous conversation you experience is reconstructed from scratch on every single request.

Understanding this changes how you think about context. It’s not a feature of the tool. It’s the only mechanism the model has. Manage it well and you get a sharp, focused collaborator. Let it fill with noise and you get drift.


What’s Actually in Every Request

Before you can manage a budget, you need to know what you’re spending. Every request Claude Code sends contains more than you probably think:

Source Typical cost Notes
System prompt ~6k tokens Always present
Tool & MCP schemas ~12–70k tokens Loaded every turn, used or not
Active skills ~3k tokens Always present
Memory / CLAUDE.md <1k–10k+ tokens Always present, every turn
Conversation history Grows per turn All prior messages, reconstructed
File reads ~3–4k per file Stays in context for the session
MCP tool results 10–15k+ per call Large structured responses can be enormous

The part that surprises most people: tool and MCP schemas are loaded on every single request, whether you use those tools or not. The API requires Claude to receive the full JSON schema of every available tool so it knows what’s available. Connect three MCP servers and you’re paying 50–70k tokens per turn just in schema overhead — before you’ve typed a word.

This means disconnecting an MCP server you’re not using isn’t just tidiness. It’s the same economics as trimming your CLAUDE.md.


Context Is a Budget, Not a Log

Most people treat the context window like a transcript — a running record of everything that happened. That framing leads to bad habits. You add things freely. You don’t remove anything. The window grows and you trust that Claude will find what it needs in the pile.

The better frame: context is a budget. You have ~200,000 tokens (up to 1M on some models). Everything above is always being spent. Some of it is worth the cost. Most isn’t.

The silent killers are file reads, MCP responses, and the always-on overhead you forget about. A single tool call returning a large JSON document can cost 10–15k tokens. Read 5 files and run 3 MCP calls and you’ve burned 50k tokens before writing a line of code. And the system prompt, tools, skills, and CLAUDE.md are sitting on top of all of it, every turn.

When you treat context as a budget, you start asking “is this worth including?” before every tool call, every file read, every CLAUDE.md addition. That question is the skill.


Bigger Window, Same Problem

A common reaction to context limits is: just get a bigger window. Claude supports up to 1M tokens — surely that makes all of this irrelevant?

It doesn’t. And not just because 1M fills faster than you expect.

The deeper problem is that model quality degrades well before the window fills. Research on large context models consistently shows that reasoning quality begins to drop noticeably somewhere around 50–60% capacity. At 600k tokens in a 1M window, you’re not getting the same Claude you had at 100k. The model can still process everything — it doesn’t error, it doesn’t warn you — but it starts missing things, losing the thread of earlier constraints, producing answers that are longer and less precise.

This is sometimes called the “lost in the middle” problem: information buried deep in a long context is attended to less reliably than information near the beginning or end. The more you’ve accumulated, the harder it is for the model to find the signal.

A 1M window is genuinely useful — it raises the ceiling for things like large codebase analysis or loading a full set of documents at once. But it doesn’t change the economics of attention. It just means you can be sloppy for longer before things go wrong. The discipline of keeping context clean matters at 100k and it matters at 500k.

More window doesn’t buy you better reasoning. It buys you more rope.


You’re the Custodian, Not Claude

Claude has no awareness of what it costs to keep things around. It can’t see how full the window is. It can’t choose to forget something. It can’t flag that a tool result from three tasks ago is still sitting in context eating attention. It reasons from whatever is in front of it — all of it, every turn — and it does so without complaint.

This means context rot is invisible from Claude’s side. The degradation shows up on your side, and it doesn’t announce itself. What you notice instead:

  • Claude starts contradicting decisions you made earlier in the conversation — constraints you set, choices you agreed on
  • It “forgets” something you established thirty messages ago, even though it’s technically still there
  • Answers get longer and vaguer. More hedging, less precision. The model is working harder to find the signal and losing confidence in what it finds
  • It starts repeating context back to you in summaries rather than acting on it

None of this throws an error. The degradation is gradual, and by the time it’s obvious you’ve usually been getting quietly worse output for a while.

That’s why context management is your job, not Claude’s. The model can’t fix what it can’t see.


Compaction Is a Last Resort, Not a Safety Net

Claude Code will automatically compact your context when it gets close to the limit. This sounds like a helpful feature. It isn’t — or at least, it isn’t something you want to rely on.

Here’s what auto-compaction actually does: when the window fills, Claude Code compresses earlier turns into a summary and discards the original content. You lose the actual decisions, the actual constraints, the actual reasoning. What replaces them is a lossy summary — a map of where you’ve been. The model now has to work from that map for the rest of the session.

The problems compound:

You don’t choose what gets summarised. Auto-compaction happens at the limit, on the system’s terms, not yours. The constraint you carefully established in turn 12 might compress down to “the user mentioned some preferences about code style.” The nuance is gone.

You often don’t notice. There’s no loud alert. The session continues. Claude keeps responding. But it’s now reasoning from a summary, and subtle things start slipping — the same kinds of symptoms as context rot, except now they’re permanent for this session.

It rewards bad habits. If auto-compaction bails you out, you have no incentive to manage context proactively. The session limps forward instead of being reset on your terms with a clean handoff.

Manual /compact is better than auto-compaction for one reason: you choose when it happens. Running it while the context is still reasonably clean means the summary is higher quality and less information is lost. But /compact is still lossy. It’s the right tool for “this session is getting long and I want to continue” — not a substitute for the discipline of keeping context lean in the first place.


The Sunk Cost Trap

People stay in degraded sessions because starting over feels like losing something — the thread, the setup, the momentum. It’s sunk cost applied to tokens. What you’re really holding onto is noise.

The same psychology shows up in CLAUDE.md. Some people wire their entire knowledge base into it — personal docs, runbooks, architecture notes, decision logs. The idea is compelling: Claude always has everything it might need, always present. The problem: Claude reads CLAUDE.md on every single turn. A 10,000-token CLAUDE.md is 10,000 tokens of overhead on every message, even when you’re asking something that has nothing to do with any of it. You keep adding because it feels like progress. The window gets heavier. The signal gets harder to find.

Both traps share the same structure. The longer a session runs, the harder it feels to cut it. The more you’ve added to CLAUDE.md, the harder it feels to trim it. The investment feels real. The cost feels abstract. But the cost is paid on every single turn, whether you notice it or not.

The question to ask is not “how much have I invested in this session?” It’s “if I were starting fresh right now, what would I actually bring?”


/rename and /resume: Keeping Sessions Findable

Before you can exit a session cleanly, it needs to be findable.

/rename gives the current session a name. This sounds minor. It isn’t. Claude Code creates sessions automatically with generated IDs. Without a name, your three-hour session on a particular bug or feature is indistinguishable from every other session you’ve ever run. When you close it and come back the next day, you have no easy way to find it.

Name sessions the moment they have a purpose. Not after an hour of work — at the start, when you know what you’re working on.

/resume lets you return to a previous named session. Rather than starting fresh and rebuilding context from scratch, you pick up the existing session — conversation history, file reads, working state — exactly where you left it.

The /rename + /resume pair is the lightweight version of the handoff workflow. Use it when:

  • You’re stepping away from a session and expect to return soon
  • You’re switching to a different task and want to keep the current thread intact
  • You want to hand a session to someone else who needs the full context

For longer pauses — overnight, across a week, handing off to another agent — the /handoff and /pickup pattern is better, because it survives context decay and produces a structured document rather than relying on a live session.


/handoff and /pickup: Ending Sessions on Your Terms

The answer to the sunk cost trap isn’t willpower — it’s a better exit. If starting over feels like losing progress, it’s because you don’t have a way to carry the relevant state forward cleanly.

/handoff solves this. Before you close a session, run it. Claude Code compresses the session into a structured document: what was done, what’s in progress, what comes next, any open decisions or blockers. The knowledge survives. The noise doesn’t.

When you’re ready to continue — in a new session, on a different machine, or handed to someone else — /pickup loads that document and resumes from where it left off. The new session starts with exactly the context it needs, nothing more.

Used together, /handoff and /pickup change the economics of starting over. A fresh session with a good handoff is not a step back. It’s a clean slate with memory. The conversation history is gone; the relevant state is not.

This is the habit worth building: before you close any session you’ll want to return to, run /handoff. Then /rename it so it’s findable. The two minutes this takes is cheap compared to reconstructing context from scratch — or worse, continuing a degraded session because starting over feels too costly.


Keeping Context Clean

The goal is to never need compaction or a messy handoff in the first place. That comes down to three practices.

Keep CLAUDE.md lean and project-specific. It should contain only what Claude genuinely needs to behave correctly in this codebase — conventions, constraints, commands, things that would take too long to re-explain every session. Everything else belongs in memory files.

Claude Code has a structured memory system at ~/.claude/projects/<project>/memory/. Instead of one massive CLAUDE.md, you write individual memory files — one per topic — with a MEMORY.md index. The index is always loaded; the individual files are loaded when referenced. Think of it as the difference between keeping every book on your desk versus knowing which shelf they’re on. The knowledge is available either way — but your desk is a lot cleaner.

Use subagents for heavy exploration. When you give Claude a large task — search the codebase, debug a slow function, summarise thirty docs — all that work accumulates in your main context. Subagents isolate it. You delegate to a subagent, it does the work in its own context window, returns only the summary. Your context gets 500 tokens heavier, not 50,000.

Disable MCP servers you’re not using. As covered above: schemas load whether you use the tools or not. A server you won’t touch this session is pure overhead. Disable it in your config, re-enable when you need it. Same economics as trimming CLAUDE.md, and just as easy to reverse.


Garbage In, Garbage Out at Inference Time

You can’t fix a bad context with a better prompt.

When the window is full of noise — stale tool results, files that were relevant three tasks ago, CLAUDE.md loaded with content that has nothing to do with the current problem — the model reasons through all of it. It doesn’t skip the irrelevant parts. It can’t. Everything in context competes for attention.

The mop metaphor in the title points at this: clean the space before you start, not after you notice the mess. By then you’ve already gotten bad output. Context hygiene is something you practice at the start and throughout — not something you reach for when things go wrong.


/rewind and /branch: Surgical Context Control

When context does go wrong mid-session, you have options beyond starting over.

/rewind steps back to a specific point in the conversation and discards everything that came after. The session continues from there — your working state, the files that were read, the decisions that were made before things went sideways — all intact. Only the drift gets cut.

Use /rewind when:

  • Claude went off in the wrong direction two or three turns ago and you want to steer differently
  • A tool call came back with a massive result that polluted the context, and you want to redo it more carefully
  • You gave a bad instruction and want to correct course without starting over

/branch is /rewind with a fork. Instead of discarding the bad path, it preserves it and opens a new branch from that same point. You can explore a different approach while keeping the original thread available to return to.

Use /branch when:

  • You want to try two different approaches to the same problem and compare them
  • You’re not sure the original direction was wrong — just want to hedge before committing
  • You’re exploring a risky change and want a safe fallback

Both commands are most useful early. Every turn you let pass after context went wrong adds more noise on top. The cost of rewinding three turns is low; the cost of rewinding fifteen is often “just /clear and start over.”


Subagents: Keeping the Mess Out Entirely

/rewind and /branch fix context after it goes wrong. Subagents prevent it from going wrong in the first place.

The pattern is simple: when you need Claude to do something expensive — search a large codebase, debug a slow function across multiple logs, summarise a set of documents — that work accumulates in your context. Every file read, every tool call, every intermediate result lands in the window and stays there. By the time you have your answer, you’ve spent tens of thousands of tokens on scaffolding you no longer need.

A subagent does the same work in its own isolated context window. It reads the files, runs the queries, reasons through the problem — and then returns only the result to your main session. Your context gets one answer, not the entire record of the investigation that produced it.

Claude Code has a built-in Explore subagent designed exactly for this. Delegate to it in the sidebar. It runs the investigation in parallel, in its own window, and hands back a summary. You stay focused on the task; the noise never enters your context.

The economics are stark. An exploration that would cost 50,000 tokens in your main session costs you perhaps 500 — the summary. That headroom stays available for actual work.

This isn’t just about heavy tasks. The habit applies to anything where you’re not sure how much context it’ll consume. When in doubt, delegate. A subagent is disposable; your main context is not.


Quick Reference

Situation Command
See what’s in your context right now /context
Starting something completely different /clear
Session getting long, want to continue /compact
Claude went off the rails /rewind
Explore two paths from the same point /branch
Name the current session /rename
Return to a previous named session /resume
Pause work to resume in a new session /handoff
Resume from a handoff document /pickup
Heavy exploration without bloating context Explore subagent