stream() / generate() call gets the
messages the client sends, and nothing else. Because nothing is shared between
calls, you can safely reuse one agent instance across concurrent runs. Fanning
out per-item reviews or classifications over a shared instance keeps every run
isolated. Configure memory on the agent to persist history across calls, and
use createAgUiHandler to stream the response back.
Memory configuration is independent of model selection, so these examples omit
model and use openai/gpt-5.4-nano.
Prerequisites
- An agent in
agents/(see Agents). - An AG-UI route (see API routes for the
createAgUiHandler("assistant")pattern). - A storage backend if you choose
conversationmemory; the default in-memory driver is fine while developing.
Choose a memory mode
Configure memory on your agent to persist messages across requests. A configured agent accumulates one shared conversation on the instance, so reuse it sequentially (a single chat thread) rather than across concurrent independent runs. For per-item fan-out, create a fresh agent per run instead. To keep the stateless default explicitly (for a single-shot agent that should never persist history), setenabled: false:
Buffer memory
Keeps the last N messages. Simple and predictable:Conversation memory
Sliding window based on token count. Drops the oldest messages when the limit is reached:Summary memory
Automatically summarizes older messages to fit more context into fewer tokens:Replay trust boundary
Input validation checks caller-supplieduser and system messages. It does not
scan assistant replay or tool results as new caller instructions. This keeps
existing conversations replayable, but message roles do not prove their origin.
Your host must load assistant replay and tool results from trusted storage or
trusted model execution and ensure they belong to the authorized conversation.
Do not forward client-supplied assistant or tool messages directly as trusted
replay. The input-validation middleware does not authenticate replay provenance.
Direct writes through getMemory().add() also cross this host-owned trust
boundary; validate untrusted input before adding it.
For violations detected in assembled provider instructions, onViolation
receives [REDACTED] as content. This keeps trusted system and historical text
out of callback-based audit logs. The violation type and reason remain available.
Custom memory transactions
When transactional input validation is enabled, your customMemory backend
must implement beginTransaction(): Promise<MemoryTransaction>. Veryfront rejects
an unsupported backend before writing input. Configured conversation, buffer,
summary, and stateless memory need no changes. Without transactional validators,
the existing custom-memory interface still works.
Import the Memory and MemoryTransaction types from veryfront/agent. Your
transaction must provide these methods:
getMessages()reads a stable snapshot plus this transaction’s staged messages.add(message)stages caller input, assistant replies, and tool results and applies your retention or summarization policy to that view. It must not publish those messages to shared storage.commit()atomically checks that the snapshot is still current and publishes the validated view. If another operation added messages, cleared history, or otherwise changed the snapshot, reject without publishing. A later attempt must take a fresh snapshot and validate again.rollback()discards only this transaction’s staged work and releases its resources. It must preserve concurrent additions and clears, including after a failedadd()orcommit().
clear() and replaying an earlier snapshot: that can resurrect deleted history
or overwrite concurrent messages. Surface storage and rollback errors instead
of reporting success.
Veryfront keeps the transaction open through every provider validation in the
turn. A later validation failure rolls back the staged turn. Commit runs after
the turn finishes validation, so concurrent storage changes can reject commit
even after the provider has produced output.
For built-in stateful memory, a direct getMemory().add() or
getMemory().clear() during an active validated turn also rejects commit.
Rollback removes the rejected turn and preserves your concurrent memory changes.
Memory projection validation checks newly merged message groups after trimming.
Unchanged retained groups keep their previous validation status.
Validated turns on one stateful agent runtime run sequentially until commit or
rollback finishes. Your delegation graph must not call back into an ancestor
runtime with an active validated turn. Veryfront rejects that cycle before it
waits on memory, so the cycle cannot block later turns. Independent concurrent
calls still wait for their turn normally.
Cancelling a queued turn stops that request without persisting its input. Later
turns still wait for the active turn to finish.
The standalone RedisMemory class does not currently implement this transaction
capability. If you connect it to transactional agent validation through a custom
adapter, that adapter must supply atomic transactions. Its existing standalone
add(), getMessages(), and clear() methods remain unchanged.
Distributed memory
Agent configuration currently supportsconversation, buffer, and
summary memory. These stores belong to one agent runtime instance. Do not use
memory: { type: "redis" }: the agent configuration schema does not wire that
type into agent() memory construction, so it is rejected at validation. The
RedisMemory class and createRedisMemory() remain available from
veryfront/agent for programmatic use. For multi-instance deployments, keep a
conversation on one runtime instance, construct a Redis-backed memory manually,
or persist and restore the conversation outside the agent.
Memory operations
Access memory programmatically in API routes:getMemoryStats() returns:
Native run events
Veryfront emits native AG-UI events for tool status, input requests, child-run status, citations, and attachments. Live SSE frames carry names such asChildRunStatusChanged; durable records use CHILD_RUN_STATUS_CHANGED. Readers
also accept the earlier Custom events. Tool-status frames can omit the tool name
or set it to null; the chat decoder preserves the status in either case. Native
event payloads reserve type, elapsedMs, and emittedAt for transport metadata.
Put application timing data in nested fields.
The public buildInvokeAgentChildRunLifecycleCustomEvent and
buildInvokeAgentChildRunProgressEvents helpers retain the { type: "CUSTOM", name, value } lifecycle shape. Their schemas and publisher callbacks keep that
contract. Veryfront converts these lifecycle events to native records when it
prepares them for durable publication.
An oversized tool-status, input-request, or child-run record becomes a
conversation-run-event-omitted marker. The marker records the original event
type and tool call ID when available; it does not retain the oversized payload.
You must keep lifecycle payloads within the run event size limit to preserve their
full content in stored history.
Streaming
Server-side streaming
UsecreateAgUiHandler() for chat UI routes. It validates the request, invokes
the agent, and returns AG-UI SSE:
agent.stream() directly only when you are building a custom transport or
non-chat streaming surface.
Persisting finished conversations
PassonComplete to persist the finalized conversation server-side after a
successful run. It is the counterpart to the client-side useConversationChat path.
It fires once, only on success, after the stream is fully flushed and closed, so
a slow or throwing persistence never delays or corrupts the response:
onComplete does not fire when the run errors or when the client
disconnects before the stream finishes. A rejected callback is caught and logged
rather than rethrown. For createAgUiRuntimeHandler, the same finalized
messages (and full response) arrive on the onFinish lifecycle context.
Client-side consumption
TheuseChat hook handles the streaming protocol automatically:
Reading run events
A run’s durable event log is available from the Veryfront API. Read it withformat=typed and every row carries a catalogued event_type, a payload named
by that type, and a span envelope (run_id, event_class, span_id,
parent_span_id, turn_id, origin_event_type, origin_custom_name,
unrecoverable_fields). No typed row uses event_type: "CUSTOM".
The veryfront/run-events module owns the reader’s half of that contract, so
you do not restate the vocabulary or the payload shapes in your own code:
RUN_EVENT_TYPES,isRunEventType, andRUN_EVENT_CLASSESfor the catalogued vocabulary, andgetRunEventClassto tell a self-containedfactfrom an order-dependentdelta.toRunEventWireNameandfromRunEventWireNameto move between a stored type such asURL_CITEDand the SSE wire nameUrlCited.getRunEventEnvelopeSchema,getTypedRunEventRowSchema, andparseTypedRunEventRowfor the row itself. Conversation-scoped surfaces (GraphQL, MCP, and the conversation events route) key the payload aseventrather thanpayload; usegetConversationTypedRunEventRowSchemathere.- One payload schema per type, such as
getUrlCitedPayloadSchema, plusRUN_EVENT_PAYLOAD_SCHEMASto look one up by type at runtime. The exception is the sixteen control-planeAGENT_RUN_*types: the API owns their shape and sanitizes it before a reader ever sees it, soRUN_EVENT_PAYLOAD_SCHEMAShas no entry for them androw.payloadis already the value to use.
SchemaValidator
contract. Inside a Veryfront app, bootstrap registers it before handlers run.
Anywhere else, including a browser bundle that reads run events directly,
register one yourself before the first get*Schema() call or
parseTypedRunEventRow:
register replaces whatever is registered, so the tryResolve gate is what
makes this safe to paste into an app that already owns a validator: it installs
the adapter only when nothing has. The module ships no fallback validator: with nothing
registered, a getter throws an error naming the SchemaValidator contract and
this registration call.
event_type is validated as a non-empty string, not as the closed catalog, so
a type the API adds after your build still parses. Narrow it with
isRunEventType when you need the closed set, and ignore what you do not
handle. Do the same for the wire names on a stream: advance the durable cursor
for every frame that carries an id, including the ones you do not render.
Eight of these types come from the Veryfront Code runtime itself:
TOOL_CALL_STATUS_CHANGED, INPUT_REQUEST_CREATED, INPUT_REQUEST_UPDATED,
CHILD_RUN_STATUS_CHANGED, URL_CITED, DOCUMENT_CITED, FILE_ATTACHED, and
RUNTIME_EVENT_RECORDED. NATIVE_RUN_EVENT_TYPES lists them.
Non-streaming generation
Usegenerate() when you need the complete response at once:
Verify it worked
Send two messages on the samethreadId (with conversation memory) and
confirm the second response references the first message. With curl: