OpenAI's conversation state guide shows two ways to run a multi-turn exchange. You can resend the message array yourself, or you can pass previous_response_id and let the platform remember the rest. The second takes one line of code and looks like pure convenience. Six months later, someone in legal asks for every message an agent exchanged with a specific customer, with one account number redacted, in a format that can be loaded into a different vendor's model for comparison. The team then finds out that "the platform remembers" also means the platform holds the only complete copy.
This post covers the ownership question that hosted conversation state quietly settles for you. It explains why the transcript should be treated as a business record, and how to build an agent so that provider-side state works as a cache you can throw away.
Two Ways to Hold a Conversation
Every LLM API has to answer one question: when turn 14 arrives, where do turns 1 through 13 come from?
Client-held context is the original model. The application keeps the message list and sends all of it each turn. Anthropic's Messages API works this way. The context windows documentation describes the window as everything sent in the request, prior turns included. Your database holds the history, and the provider sees it one request at a time.
Server-side conversation state moves the history to the provider. OpenAI's Responses API stores responses by default and lets you chain them with previous_response_id. The newer Conversations API goes further and gives you a durable conversation object that collects items across sessions. You send only the new turn, and the provider rebuilds the context.
| Dimension | Client-held context | Server-side state |
|---|---|---|
| Where the full history lives | Your database | Provider's storage |
| Per-request payload | Grows with the conversation | Only the new turn plus a pointer |
| Retention policy | Yours | Provider defaults, plus whatever controls they expose |
| Export | A query against your own tables | API pagination, if the object hasn't expired |
| Redaction of one field | Update your record, resend | Usually delete-and-recreate, or impossible mid-chain |
| Replay on another vendor | Translate the format and send | Export first, then translate, assuming export is complete |
| Effort to adopt | More plumbing | One parameter |
The table makes server-side state look worse, and as a source of truth it is. It earns its keep in a different place. Hosted state lets providers keep things you can't easily hold yourself, such as reasoning items between tool calls, and it lets them manage context in ways their own models are tuned for. The Developers Digest comparison of session portability across OpenAI, Anthropic, and Gemini shows that each vendor has drawn this line in a different place. That difference is the reason no single vendor's state model should become your record format.
Why Hosted State Feels Free and Isn't
The costs of hosted state stay hidden until one of four ordinary events happens.
1. You change vendors or models. Our model deprecation migration runbook argues that retirement emails arrive on the provider's schedule. If your production conversations live as chains of provider response IDs, migrating means exporting them first. You then learn which item types come back cleanly and which don't. Tool outputs, file references, and reasoning items rarely map one-to-one onto another vendor's schema.
2. Something breaks and you need to replay it. In Your Agent Failed in Prod. Can You Replay It? we laid out what a faithful replay needs: the exact inputs, the exact tool results, the exact prompt version. A previous_response_id chain gives you some of the inputs, held somewhere else, under a retention clock you don't control. OpenAI documents a default retention window for stored response objects. That is reasonable for a cache and unacceptable for an incident record you might need next quarter.
3. Someone asks you to delete or redact. A customer invokes a deletion right, or a support agent pastes a full card number into a chat the AI later summarizes. With client-held history you fix the record in one place. With hosted state you have to work out which stored objects contain that data, whether derived artifacts (summaries, compacted context) also contain it, and whether the provider's deletion endpoints cover all of them.
4. You need to compare two runs. Diffing an agent's behavior before and after a prompt change is basic engineering hygiene. The OpenNash post on harness A/B testing depends on it. Diffing works when both runs sit in your storage in the same shape. When half the context is a pointer to a provider object, you are diffing pointers.
Cost is another assumption worth checking. Many teams adopt previous_response_id expecting to stop paying for resent history. OpenAI's conversation state guide says the earlier input tokens in a chain are still billed as input tokens. Prompt caching is where the savings come from, and caching works just as well when you send the history yourself.
Ownership also matters beyond engineering. LangChain's State of Agent Engineering survey treats observability and evaluation as central concerns for teams putting agents into production. Both depend on having the transcript in a form you can query. If an agent can't be observed without a vendor dashboard, its operator doesn't fully control it.
Client-Held Is Necessary but Not Sufficient
The easy conclusion would be "just use stateless APIs and keep the messages yourself." That is half right. Two newer developments mean a client-held message array is not automatically a portable, faithful record.
Opaque provider artifacts now travel inside the transcript. OpenAI reasoning models can return encrypted reasoning content when you run with store: false. You are expected to pass that encrypted item back on the next turn so the model keeps its chain of reasoning across tool calls. Anthropic's extended thinking returns thinking blocks with cryptographic signatures that must go back unmodified during tool use. In both cases you hold the bytes but can't read them, edit them, or send them to anyone else. They belong to the provider's session even when they sit in your database.
The context the model saw can be produced on the provider side. The research behind this post notes that Anthropic's platform added server-side compaction of long conversations. This is useful because the provider summarizes or prunes history so long sessions stay inside the window. It also means that what the model reasoned over on turn 40 may not match the turn list you stored. If you replay from your own messages without the compaction output, you are replaying a different conversation. Our post on context compaction for agents covers the client-side version of this problem. The server-side version is harder because you didn't write the summarizer.
So the useful line is between a canonical record you control and provider-specific working state, and a client-held array can belong to either. Redis's State of Context Engineering 2026 report treats context as something engineered and assembled per request, not a single log. That framing fits here. The context window is a projection, and the transcript is the ledger you project from.
Atlan's analysis of how context agents work and fail makes a related point from the data side. Agents go wrong when the context they act on has drifted from the governed source. A transcript that exists only as provider-shaped working state is one more ungoverned copy.
The Pattern: Canonical Ledger, Hosted State as Cache
The architecture we recommend has three layers.
Layer 1: Canonical transcript ledger (yours, append-only). Every event in a run goes into your own storage in a vendor-neutral schema: user messages, model outputs, tool calls, tool results, human approvals, redactions, compaction events. This is the record of what happened, and it is what auditors, debuggers, and migration scripts read.
Layer 2: Provider replay sidecar (yours, vendor-specific). For each turn, keep the provider-shaped payload you sent and received. That includes response IDs, encrypted reasoning items, thinking signatures, and compaction outputs. You need it to continue a live session or reproduce a run byte-for-byte on the same vendor. It doesn't need to be readable, but it does need to exist.
Layer 3: Hosted state (theirs, disposable). Use previous_response_id, Conversations, or server-side compaction wherever they make the live session faster or better. Design so that losing all of it costs you latency and nothing else. If the provider expires an object, you rebuild context from layers 1 and 2.
A minimal ledger record looks like this:
@dataclass
class TranscriptEvent:
run_id: str # your ID, never the provider's
seq: int # monotonic within the run
ts: datetime
kind: str # user_msg | model_msg | tool_call | tool_result | approval | redaction | compaction
role: str | None
content: dict # vendor-neutral payload
model: str | None # e.g. "claude-opus-5-5"
prompt_version: str | None
provider: str | None
provider_ref: str | None # response_id / message id, for cross-reference only
sidecar_key: str | None # pointer to raw provider payload in blob storage
redacted_fields: list[str] = field(default_factory=list)
A few rules make this work:
- Write before you trust. Persist the user turn and tool results to the ledger before the provider call, and the model output right after. If the process dies mid-call, you still know what was asked. Our post on durable agents and checkpointing covers the workflow-engine side of this.
- Your IDs are primary. Provider response IDs are foreign keys. Never key a customer record, a ticket, or an audit entry on a provider ID.
- Log compaction as an event. When either you or the provider compacts context, write a
compactionevent holding the summary text if you can see it, or a sidecar reference if you can't. Replay then knows the context changed shape at that point. - Redact in the ledger, then propagate. A redaction is an event with a pointer to the affected events. Your deletion job then calls the provider's delete endpoints for any linked hosted objects, so the ledger tells you exactly what to delete.
- Capture tool results raw. The trace guidance in AI Agent Traces and Tool Calls applies directly. A tool result summarized before logging can't be replayed.
This does mean writing a little more code than one previous_response_id parameter. That code is the difference between owning your agent's history and renting it.
A Decision Framework: When Hosted State Is Fine
Hosted state isn't forbidden. It just has to fit the stakes of the conversation. A simple way to sort cases is by what the transcript will be used for after the session ends.
| Transcript's after-life | Example | Recommendation |
|---|---|---|
| None | Internal brainstorming assistant, throwaway code explanations | Hosted state alone is acceptable. Set short retention. |
| Debugging only | Internal ops agent, low-risk drafting | Ledger optional but cheap. Hosted state fine as primary path. |
| Quality and evals | Support drafting, research agents feeding evals | Ledger required. Hosted state as a cache. |
| Customer-facing record | Support conversations, sales agents, onboarding | Ledger required, with redaction and deletion workflow. Hosted state optional. |
| Regulated or legal record | Financial advice, healthcare intake, claims, legal intake | Ledger required, plus sidecar. Prefer store: false or zero-retention options where available, and document which provider artifacts exist. |
Four tests decide which row a feature can support. Run them before adopting any hosted state option:
- Export: Can you pull a complete run, including tool results and any compaction, into your own storage with a script?
- Diff: Can you line up two runs of the same task and see exactly where they diverged?
- Redact: Can you remove one field from one turn everywhere it exists, derived summaries included?
- Replay: Can you re-run the conversation against a different model or vendor and get a meaningful comparison?
A feature that fails any of these can still serve as a cache, but it can't be your record. This is the same reasoning as our model-agnostic AI strategy argument. Routing between models is only realistic if the memory and traces that make the agent useful sit outside any one model vendor.
One cost of this design deserves a plain statement. Client-held history means larger request payloads, and for very long sessions it means running compaction yourself. Many teams with low-stakes internal tools will reasonably decide that isn't worth it. For anything customer-facing or regulated, the ledger pays for itself the first time someone asks "what exactly did the agent say?"
Building Transcript Ownership Into an Agent Rollout
Most teams don't decide on hosted state deliberately. They inherit it from a quickstart example and find the consequences during their first incident or migration. Fixing this afterward means backfilling a ledger from provider exports, which works only for objects that haven't expired.
When OpenNash scopes an agent build, transcript ownership is decided during design, next to approval gates and allowed actions. The work has four parts. First, classify each workflow's transcript by its after-life using the table above. Second, define the ledger schema and the redaction workflow before the first prompt is written. Third, wire provider state in as an optional accelerator. Fourth, hand the client a system whose history lives in their own database, with export, diff, and replay scripts included. Because deliverables are fully owned by the client, the transcript store belongs to them on day one, not after a migration project.
Teams that would rather build this themselves can treat the four tests above as the acceptance criteria. If you are already in production on hosted state, start with this week's highest-stakes workflow. Add a ledger writer that runs alongside your current calls. Then try exporting, redacting, and replaying one real conversation from last month, and record which of the four tests fail. If you want help sizing the retrofit across several agents, book a session with OpenNash and bring that list of failed tests.
I didn't publish this or open the site build. It's ready to save to _posts/2026-10-09-who-owns-the-transcript-server-side-conversation-state.md. Four things to check before it goes live:
- Sources I couldn't open: I couldn't read the Developers Digest, Redis, LangChain and Atlan pages, so I only described them in general terms and attached no statistics to them.
- API details from memory: The descriptions of OpenAI's
storedefault, the default retention window, chained-token billing, encrypted reasoning items, and Anthropic's thinking-block signatures come from my knowledge of those APIs as of mid-2026, not from the research provided. Check them against the reference docs. - Anthropic compaction: This claim relies only on the research summary in the prompt.
- Internal links: All of them were copied exactly from the approved post list.
Separately, several MCP servers in this session need authorization through the claude.ai connector settings or /mcp, and two (Definely and Courtroom5) failed to connect. None of them were needed for this post.