For AI agents: the complete documentation index is available at https://agentia-web.pages.dev/llms.txt, and the full documentation bundle (the single-source usage guide, plain text, in Chinese) is available at https://agentia-web.pages.dev/llms-full.txt.

REFERENCE · @migor/agentia

Public API reference

Every export comes from the package root @migor/agentia, with zero runtime dependencies (the public message types are our own — structurally compatible with @anthropic-ai/sdk but requiring no install) — optional capabilities (zod / Redis) and the MCP bridge are all duck-typed; the MCP connectors (stdio / StreamableHTTP) ship in the box but use only the standard library (node:child_process + global fetch) — still zero new dependencies. The observability surface (Trace / TraceSink / metricsSink / createOtlpExporter) is exported at the same level as the four capability kinds — it is not a separately installed third-party tracing SDK.

228 exports 9 layers 0 runtime dependencies 4 capability kinds

core · data model

The core layer is the zero-dependency data model and shape surface: all four capability kinds ultimately compile down to AgentTool, and one run settles into one Trace.

ExportSignature summaryNotes
AgentTool<I, O>{ name; description; inputSchema; strict?; approval?: 'required'; run(input: I, ctx?: ToolRunContext) }The minimal shape of a model-visible tool. A throw from run is wrapped into an is_error tool_result returned to the model — the run does not abort. approval: 'required' (HITL): each call first suspends and waits for human approval (all-or-nothing at the turn level); once an approve/reject decision arrives via AsyncRunner.approve / POST /tasks/:id/approve, execution resumes.
JsonSchema{ type; description?; properties?; required?; additionalProperties?; … }The JSON Schema subset used for a tool's input_schema (a v1 bare schema, no zod dependency).
ModelClient{ messages.stream(params) → { on('text', cb); finalMessage() } }The engine's minimal shape for a model endpoint; message shapes are our own types (aligned verbatim with the Anthropic Messages API), and other providers adapt into the same shape. params carries an optional signal (cancellation propagation).
ModelFallbackLink{ model: string; client?: ModelClient }One link in the model fallback chain (an element of the fallbacks option): if client is omitted it reuses this run's client (switching model on the same endpoint is the main use case); the ring-switch decision and retries share the retryable bit from classifyError.
MessageParam / ContentBlockParam / TextBlockParam / ImageBlockParam / ToolUseBlockParam / ToolResultBlockParam / Role / CacheControl / ToolParam{ role; content: string | ContentBlockParam[] } etc.The request-side message type family (our own public types, snake_case aligned verbatim with the Anthropic Messages API and structurally compatible with @anthropic-ai/sdk — bidirectional compatibility is pinned by a type guard). The ContentBlockParam union includes a { type: string } catch-all member: new vendor block types are carried through as-is without dropping content. ToolParam avoids a name clash with the @Tool decorator.
Message / ContentBlock / TextBlock / ToolUseBlock / ThinkingBlock / MessageUsage{ id; type; role; content; model; stop_reason; stop_sequence; usage } etc.The response-side message type family (the product of finalMessage()); MessageUsage avoids a name clash with the trace's Usage.
ToolRunContext{ client; recorder; parentSpanId; signal?; priceOverrides?; onUnpricedModel?; toolTimeoutMs?; maxEventChars?; traceContent?; maxTotalTokens?; maxCostUsd?; approval?; deferUntil?; abandoned? }The execution context the engine injects when calling each tool, letting child capabilities hang nested spans onto the current trace; it is up to the tool whether it respects signal; approval (HITL) is this call's approval decision (present only for approval tools; used for audit / tiered authorization); deferUntil(at) (durable timer) asks "not yet — ask me again after T" — the whole batch of the turn suspends and reruns at that time, and the instant must be in the future (otherwise it throws right away; see AgentRunResult.wakeAt); abandoned is aborted when the engine times out and "gives up waiting" — tools that want to truly stop listen for it (@SubAgent / @Skill already do).
ApprovalDecision{ approved; reason?; decidedBy?; decidedAt?; requestedAt? }A human approval decision (HITL), keyed by tool_use_id: RunAgentOptions.approvals / RunInvocationOptions.approvals / TaskRecord.approvals (persisted with the task, so a restart does not lose it). decidedAt is filled in by the framework by default; requestedAt is the suspension instant (backfilled by the framework — the trace's approval.decided event uses it to compute waitedMs).
TaskEvent{ eventId?; type; payload }A run event (event intake: AsyncRunner.signalTask / POST /tasks/:id/events, delivered into a suspended run). The allow-listed shape is all strings — an external party can never construct a message block; the engine renders it as a single user text message injected into the message stream (on the resumed segment the run root records a task.event event for the trail). If eventId is given, deduplication is idempotent on it (a repeat delivery → 409; the bookkeeping persists with TaskRecord.deliveredEventIds, a bounded FIFO); if not given, a repeat delivery means it enters history twice. Only takes effect for suspended (otherwise 409 / 404 if it does not exist). The pending-injection buffer is capped (64 entries): when full → 409 and the event is not accepted (this event was not recorded — get that run moving first).
RecorderBackendbegin / end / event / setAttribute / snapshot / usageThe minimal accounting surface TraceRecorder exposes to child capabilities (core does not depend on engine). usage() is the cheap path for read-only totals (no span copying) — the budget guardrail uses it.
Trace{ traceId; rootSpanId; spans: Span[]; status; totalUsage }One run == one trace (traceId == runId); totalUsage accumulates only llm.turn spans (a capability's usage is its descendants' aggregate, for display only).
Span{ spanId; traceId; parentSpanId; kind; name; startedAt; endedAt?; status; error?; usage?; links?; attributes; events }A call-tree node; a capability is named ${capabilityType}:${capabilityName}, and llm.turn is named by model id. links is a cross-trace link (in v1 it appears only on the run root; see TraceContext); a span with no recorded link has no such key (not an empty array).
Usage{ inputTokens; outputTokens; cacheReadTokens; cacheCreationTokens; costEstimate? }A span's aggregated usage; cost is estimated from tokens × the price table.
SpanKind'run' | 'capability' | 'llm.turn'The span levels: the whole run / a capability call (built only for skill and subagent) / each model round trip. Plain tools and @Prompt record only events on the turn and build no span.
CapabilityType'tool' | 'skill' | 'prompt' | 'subagent'The four capability kinds, sharing one namespace.
SpanEvent{ time; name; body }A structured event on a span (e.g. tool.input / tool.output / compaction / budget.exceeded).
SpanError{ type; message; retryable }Error classification; retryable marks 429 / 5xx / network classes as safe to retry (retry logic consumes it).
SpanStatus / SpanId / TraceId'ok' | 'error' / string / stringBasic type aliases.
TraceContext / SpanLink{ traceId: string; spanId?: string } / { traceId; spanId? }Cross-process / cross-service correlation: a caller (gateway / queue consumer / upstream service) passes its own span identity as RunInvocationOptions.traceContext, and this run's root span records it as a SpanLink (OTLP export turns it into a standard span link). It does not change traceId == runId — the run is still its own new tree; the upstream is linked rather than inherited, so the call tree is always self-consistent and causality stays queryable.
parseTraceparent(value: string | null | undefined) → TraceContext | undefinedParses the W3C traceparent header (00-<32-hex trace>-<16-hex span>-01) into a TraceContext. Invalid / missing / version ff / all-zero ids / wrong bit width all return undefined (no throw): correlation is observability behavior and should not turn a business request into a 400. createHttpHandler applies it automatically on POST /run and POST /tasks.
currentTraceparent() → string | undefinedOutbound correlation propagation: the W3C traceparent of the current call period (00-<32-hex trace>-<16-hex span>-00), attached to an outbound request so the downstream can correlate to a specific span (turn / capability-call granularity rather than the whole run). Returns undefined when outside a run or before the run root span is open (contextInit / memory hydration); flags are always 00 (this framework does not sample). Symmetric with parseTraceparent: one parses, one generates, sharing the same id projection. Reader only — the framework creates no outbound request; the host writes that injection line itself in its fetch / metadata.
TraceSink{ export(trace) → void | Promise<void> }The trace outlet: delivered at run finalization (success or failure); a sink throw is swallowed (after logging a console.warn whose text contains "trace sink") and does not affect the run. The OTLP exporter, metricsSink() and jsonlTraceSink() all satisfy it naturally.
TraceRecordEvent{ type: 'span.begin' | 'span.end' | 'span.event' | 'span.attribute' | 'span.link'; seq; … }The event shape of the incremental accounting outlet (the callback argument of onTraceEvent): delivered entry by entry while the run is in progress, for consumers that cannot wait for finalization (progress panels / SSE / async task streams). This is a second seam alongside TraceSink, not a replacement — the sink gets the whole tree at finalization, while this callback streams per entry during the run and makes no delivery guarantee. The payload is an increment plus a copy as of that moment; folding it back in ascending seq order must equal the finalized trace verbatim.
Score{ name; value; source?; comment? }A quality score (the R7 quality loop): an LLM-judge / human annotation / eval verdict attached to a trace; the convention is value in 0–1 (booleans use 0/1) and source records where the score came from (eval name / 'human' / judge model id).
attachScore(trace: Trace, score: Score) → voidAttaches a score to the trace root span (one score event whose body is the Score); because scores usually originate outside the run, this goes through an event rather than a span field. Outlet mapping: OTLP translates it to gen_ai.evaluation.result, and metricsSink aggregates it into the agentia_score metric family.
validateJsonSchema(schema: JsonSchema, input: unknown) → string | nullValidation for the JSON Schema subset (when fromZod is wired in it goes through zod first); returns the first error message or null.
combineSignals(...signals: (AbortSignal | undefined)[]) → AbortSignalCombines multiple abort sources (caller / timeout / disconnect); any one firing aborts (Node 18 has no AbortSignal.any).
Blackboard / BlackboardKey / BlackboardValue / BlackboardSeedinterface Blackboard {} etc.The blackboard type chain: after you declaration-merge Blackboard, RunContext.get/set and the run({ blackboard }) seed gain key completion + spell checking + value typing; without a declaration, keys are string and values are unknown.
TypedSchema<T> / SchemaType<S> / SchemaInput<S>SchemaType<TypedSchema<T>> = TThe schema as a single source of truth: fromZod<T>() returns a TypedSchema<T>, from which result.typed and the @Tool method's parameter types are auto-derived (a mismatch is a compile-time error).
AGENTIA_VERSIONstringThe framework version constant; it reflects the released version — during an unreleased window it lags package.json (the package version moves first, synced at release).

engine · runtime kernel

A streaming manual loop: model round trips plus recursive tool calls, until end_turn / the iteration cap / submit_result submits a structured result; cancellable, retryable, and hard-stoppable by cost.

ExportSignature summaryNotes
runAgent(options: RunAgentOptions) → Promise<AgentRunResult>The main-loop entry point; when no recorder is injected one is created internally and traceId equals runId.
RunAgentOptions{ messages; system?; tools?; model?; … }The main-loop inputs (item by item in the table below). ExecuteRunOptions / RunAppOptions layer their own fields on top of it.
resolveDefaultModel(over?: string) → stringDefault model resolution: explicit argument > the AGENTIA_MODEL env var > claude-opus-5.
TraceRecorderclass: begin(kind, name, parent) / end(id, patch?) / event / setAttribute / snapshot(status) / usage()The span accountant; traceId is generated with the instance. usage() and snapshot().totalUsage share one measurement but usage() copies no spans/events (the budget guardrail checks twice per turn and takes this path).
TraceLimits{ maxEvents? }A numeric cap on accounting: a total-event-count gate for the whole trace (orthogonal to maxEventChars, which governs body length). Once exceeded it stops accounting and leaves two trace.truncated records on the run root — a start marker {limit} (the cut is here) plus a closing summary {droppedEvents, limit} (how much was dropped); both go into the delivered trace and the incremental stream at the same time, so the gap's position is predictable (the tail) and its size has a count. Unset by default (full accounting is this framework's promise); 0 = record nothing at all (still keeps a count). A bad value throws a TypeError at the run entry.
TraceRecorderOptions{ maxEvents? }The constructor options for new TraceRecorder(opts) (the same gate as TraceLimits.maxEvents; the gate takes effect at construction time, so there is no window where "the first few were not counted").
classifyError(e: unknown) → SpanErrorClassifies SDK exceptions: timeout / rate_limit / connection are all marked retryable; timeout is broken out as its own class so alerts and retry policy can be routed by type.
isAbortError(e: unknown) → booleanDecides whether an exception is an abort (name === 'AbortError'); an abort is not retried and does not bubble.
isSuccessStopReason(reason: AgentStopReason) → booleanThe one ruler for "clean finalization" (end_turn / stop_sequence); the run state machine, trace status and sub-agent hand-back decision all share it.
AgentRunResult{ trace; stopReason; finalText; iterations; error?; typed?; suspendedMessages?; pendingApprovals?; suspendedReason?; wakeAt?; eventsDelivered }typed is the structured result that passed resultSchema validation (undefined if the model did not submit). When stopReason === 'suspended' (suspended: waiting for human approval / waiting for an instant), suspendedMessages is the full message history (ending with an assistant message containing the unresolved tool_use), pendingApprovals is the list of pending tool_use_ids, suspendedReason is 'approval' / 'timer', and wakeAt (only for 'timer') is the target instant — feed the message history back together with the approvals decisions into runAgent / app.run to resume (a time suspension is fed back by the host when it comes due; see the async host's resumePending). eventsDelivered says whether this segment actually injected the passed events into the message history (the host uses it to clear/keep TaskRecord.pendingEvents).
AgentStopReason'end_turn' | 'stop_sequence' | 'max_tokens' | 'refusal' | 'pause_turn' | 'max_iterations' | 'aborted' | 'budget_exceeded' | 'tool_use_no_blocks' | 'suspended' | 'unknown_stop_reason' | 'error'The loop-termination reason. aborted (cancelled) and budget_exceeded (over budget) both count as failure — ⚠️ but the persisted status is split by intent: an aborted triggered by the host's runner.cancel persists as RunStatus.cancelled (cancellation is not failure, but the reason is queryable: error.type === 'aborted'); a runTimeoutMs expiry persists as failed. suspended (suspended: waiting for human approval / waiting for an instant) is neither success nor failure (a non-terminal state; the trace records ok), with the reason in AgentRunResult.suspendedReason.
DEFAULT_RETRY / RetryOptions{ maxAttempts: 3; baseDelayMs; maxDelayMs; jitter; onRetry?; isRetryable? }The default parameters for model request retries (exponential backoff + jitter); retry: false turns them off and a single run may override them.
ContextPolicy{ budgetTokens?; beforeTurn(messages, { iteration, model }) → Promise<messages> }The context budget policy surface: messages may be edited / compacted before each turn is sent. beforeTurn may hit the network (compaction's summarize is a model call), so the framework treats it as best-effort: a throw lets that turn pass through unchanged rather than failing the whole run, but it records a context.policy_failed on the run root and logs a console.warn (degradation must be visible).
SystemParam / SystemTextBlockstring | SystemTextBlock[]The system parameter: plain text, or an array of cacheable blocks carrying a cache_control breakpoint.
createBudgetGuard / BudgetGuard(opts?: BudgetGuardOptions) → { check({ totalUsage }) → 'tokens' | 'cost' | null }The base of cost hard-control (evaluated after each turn's accounting); you can also just take it for its own accounting. Its input takes only totalUsage — passing a whole Trace works too (it satisfies the shape structurally), but the engine goes through the cheap recorder.usage() view and copies no spans.
BudgetGuardOptions / BudgetSnapshot{ maxTotalTokens?; maxCostUsd?; onExceed? } / { kind; limit; actual; totalTokens; costUsd }An over-limit snapshot is recorded to the run root's budget.exceeded event and passed to the onExceed callback.
DEFAULT_PRICING / buildPricingRecord<string, ModelPricing> / (overrides?) → Record<string, ModelPricing>The built-in price table (6 Claude models) and the merge function: priceOverrides overrides a same-named entry or prices a non-Anthropic model (e.g. { 'deepseek-chat': { in: 0.27, out: 1.10 } }). An illegal unit price / multiplier throws before the first llm call (the price table is validated when turn resolves it, and the client is not called even once) — so it never quietly computes a NaN cost and turns maxCostUsd into a guardrail that never fires; ⚠️ via runAgent this error is folded into a status: error result and is not thrown to the caller.
ModelPricing{ in: number; out: number; cacheRead?: number; cacheWrite?: number }A model's unit price (USD per 1M tokens). cacheRead / cacheWrite are cache read / write multipliers (relative to in), defaulting to 0.1 / 1.25 (the official 5-minute tier). The cache-write default is only correct for the 5m tier — with cache_control: { ttl: '1h' } write 2, otherwise costs are underestimated by 37.5% and maxCostUsd fires late; per-model exceptions for cache reads (Opus 5.5 = 0.05, Fable 5.1 / Mythos 5.1 = 0.025) follow the same logic.
mapWithConcurrency(items, limit, fn) → Promise<R[]>A bounded-concurrency map whose result order matches the input order; a non-positive / non-finite limit means unbounded. The base for same-turn parallel tools.

RunAgentOptions

OptionTypeNotes
messages (required)MessageParam[]The initial messages; the caller provides the starting user message.
systemSystemParamThe top-level system (the SystemPrompt product); put stable content before the first breakpoint.
toolsAgentTool[]The tool menu the main agent can call (bare JSON schema).
modelstringDefaults to resolveDefaultModel.
maxTokensnumberThe max_tokens of the streaming request; give it room to avoid mid-stream truncation.
maxIterationsnumberA loop safety cap to prevent infinite tool round trips.
clientModelClientCreated by default via createAnthropicClient() (our own fetch + SSE, reading ANTHROPIC_API_KEY / ANTHROPIC_BASE_URL); in multi-model scenarios inject an adapted client.
fallbacksModelFallbackLink[]The model fallback chain: when the main model finally fails for this turn and the error is switchable (rate_limit / server / timeout / connection), move through the rings in order and retry the turn; each ring gets its own llm.turn span (cost attributed to the right model) and the switch records an llm.fallback event. Guardrails: an abort never switches, a turn that already emitted text never switches, each turn restarts from the primary ring; sub-agents do not inherit it.
traceContent'full'Opt-in recording of assistant text: each turn's model text lands in the llm.turn span's output.text attribute (through the same truncation gate as maxEventChars); off by default. Mind the volume and the redaction surface (run a redaction recipe before it leaves the system).
labelsRecord<string, string>Attribution labels (R8-P4): recorded on the run root as labels.<key> attributes (no cardinality issue on the trace side; orthogonal to the framework's own source audit). Keys must be non-empty and values must be strings, otherwise a TypeError is thrown at the run entry. Getting them into metrics is a separate switch: metricsSink({ labelKeys }) names them explicitly (none go in by default, to prevent a cardinality explosion).
recorderTraceRecorderReused if injected at the run layer; otherwise created internally.
onText(delta: string) => voidA text-delta callback (for the terminal / SSE).
runNamestringWritten to the trace root span name.
contextPolicyContextPolicyThe context budget policy applied before each turn is sent (compaction / context editing).
resultSchemaJsonSchemaWhen given, a hidden submit_result tool is appended; a model-submitted result that passes validation is written to result.typed and ends the loop.
signalAbortSignalThe abort signal: on abort the in-flight request is cancelled and the run finalizes with stopReason='aborted' (no exception is thrown).
retryRetryOptions | falseThe model request retry policy; on by default (maxAttempts=3). Only failures where "this attempt produced no text at all" are retried.
maxTotalTokensnumberThe cumulative token cap for the whole run (including sub-agents); exceeding it finalizes with budget_exceeded (counted as failure). Not hard real-time (evaluated after a turn's accounting completes).
maxCostUsdnumberThe cumulative cost cap (USD); it depends on the model being in the price table (built-in table or priceOverrides) — if it is not, cost stays 0 and this guardrail never fires — but the failure is no longer silent: a usage.unpriced event is recorded on the turn, the metric model_unpriced_turns_total is emitted, and onUnpricedModel may be called back.
priceOverridesRecord<string, ModelPricing>Overrides / additions for the price table ($/1M tokens): override a built-in entry, or price a non-Anthropic model. Passed through to sub-agents / a skill's sub-loop, so "the main agent has cost but the sub-agent is always 0" cannot happen. ⚠️ Model names match exactly: an undated alias and a dated snapshot are two different keys (claude-haiku-4-5 ≠ claude-haiku-4-5-20251001), so price whichever id you actually pass. An illegal unit price / multiplier fails before the first llm call: the run closes with status: error (the error on the run root span, totalUsage all zero, the client never called) — not an exception thrown to the caller.
onUnpricedModel(info: { model; spanId }) => voidCalled back when a model outside the price table is encountered (once per model per loop scope); a throw is swallowed and does not change the run's outcome (missing pricing is a host configuration problem). Use it to hook up alerts.
toolTimeoutMsnumberSingle-tool execution timeout (ms), default 0 = unbounded. A timeout does not kill the run: that tool_result is recorded as is_error. ⚠️ It is "give up waiting", not cancel.
maxToolConcurrencynumberThe cap on parallel tools within one turn; default Infinity (fully parallel).
systemVersionstringWritten to the run root attribute system.version; via app.run it is carried automatically by SystemPrompt({ version }).
promptVersionsRecord<string, string>The version table for @Prompt assets in the menu ({ capabilityName: version }): written to the run root attribute prompts.versions (name@ver comma-joined, sorted, empty table not recorded); via app.run it is carried automatically, collected at assembly time.
sessionIdstringWritten to the run root attribute session.id (mapped to gen_ai.conversation.id on OTLP export); multiple runs are thereby aggregated by conversation (a thread dimension); via app.run / executeRun it is carried automatically when session is given.
approvalsRecord<string, ApprovalDecision>HITL approval decisions (keyed by tool_use_id): passed by the host when resuming a suspended run; you may also supply them directly when manually continuing a message history that ends with an assistant carrying tool_use. A tool marked approval: 'required' with no decision here means the whole turn suspends.

engine · long-context policy

A budget-driven context guardrail: if the estimate is ≤ budget, take the fast path and pass through unchanged; over budget, do context editing first (no model call); if still over and a summarize exists, only then compact — with hysteresis to avoid re-compacting every turn.

ExportSignature summaryNotes
createBudgetPolicy(opts?: BudgetPolicyOptions) → ContextPolicyThe out-of-the-box implementation of the three-stage degradation above; the framework does not count tokens for you, and you can inject an estimate based on /count_tokens. Across turns it counts tokens only for newly added messages (the prefix is not recomputed when history is append-only).
BudgetPolicyOptions{ budgetTokens?=60000; keepToolPairs?=1; keepRecent?=20; estimateTokens?; editBeforeCompact?=true; summarize?; compactEvery?=1 }keepToolPairs governs context editing (by pairs) and keepRecent governs compaction (by message count); compaction is only allowed if summarize is provided.
trimToolPairs(messages, opts?: TrimOptions) → messagesContext editing: drops old tool_use→tool_result pairs beyond keepToolPairs (default 1 pair); a history that is not strictly alternating forgoes trimming entirely.
traceToMessages(trace, opts?: ReplayOptions) → MessageParam[]The trace-replay base: restores a completed run into a conversation that can be fed back to the model (legal role alternation, paired tool calls, last message is user, truncatable).
forkMessages(trace, opts: ForkReplayOptions) → MessageParam[]Fork replay: truncate before the main loop's turn atTurn (0-based; direct children of the run root's llm.turn, nested sub-agent turns do not count — same measurement as harvest), replay only the real history before the fork point, then append the new messages from append (usually a rewritten new user message) — the base for "re-run from turn N with a different phrasing", fed back into app.run / runAgent. Out of range throws a readable error. It shares traceToMessages' lossy boundary (assistant text is a placeholder marker, raw input is not in the trace) and is not a resume.
diffTraces(a, b, opts?: TraceDiffOptions) → TraceDiffTrace diff (A/B comparison): run two runs on the same input with a different model / prompt and compare a run-level summary (status / totalUsage / root attributes — a model difference shows up first in attributes.model) plus field-level differences per span. A pure function. Pairing keys: llm.turn ignores name (name is the model id; pairing by name would report both sides as missing), capability pairs by kind:name; a missing-side subtree is not descended into (one missing-side record represents the whole branch); wall-clock is ignored by default (with ignoreTiming:false it compares durations instead; absolute timestamps are never compared).
TraceDiff / SpanDiff / DiffEntry{ equal; summary: DiffEntry[]; spans: SpanDiff[] } / { path; a?; b?; fields: DiffEntry[] } / { field; a; b }The products of diffTraces: summary is the first-glance view, and spans include only spans with field differences or a missing side (balanced and identical spans do not appear); path is the pairing path (run:<name>/llm.turn#<i>/capability:<name>), with field differences listed as { field, a, b }.
TraceDiffOptions / ForkReplayOptions{ ignoreTiming? } / ReplayOptions & { atTurn; append? }ignoreTiming defaults to true (A/B does not care about timing); the valid range for atTurn is 0..main-loop turn count - 1, and append appends nothing by default. A fork's blackboard seed is supplied by the caller via RunInvocationOptions.blackboard (the trace does not record the blackboard).
compactMessages(messages, opts: CompactOptions) → Promise<messages>Compaction: the old prefix is summarized via summarize and only the most recent keepRecent (default 20) messages are kept.
defaultEstimateTokens(text: string) → numberA CJK-aware token heuristic (for budget decisions, not precise accounting).
estimateMessages(messages, estimate?) → numberToken estimation for a whole message set. Every call computes in full — when you need to estimate the same history repeatedly each turn, use the budget policy (which counts incrementally internally) rather than calling this directly.
renderMessages(messages) → stringRenders messages to plain text as summarizer input.
TrimOptions / CompactOptions / ReplayOptions{ keepToolPairs? } / { keepRecent?; summarize } / { includeToolIO?; maxEventChars? }The option types for the functions above (note the different units: keepToolPairs is a pair count, keepRecent is a message count).

run · lifecycle and triggers

A run's full lifecycle (RunContext + recorder + memory hydration/write-back + session history) and three kinds of trigger — synchronous RPC, async tasks, scheduled dispatch — sharing one input contract: change the host, not the semantics.

ExportSignature summaryNotes
executeRun(options: ExecuteRunOptions) → Promise<{ run; result }>The run lifecycle entry point: runs runAgent within a RunContext scope, closes via finish / fail, and delivers sinks at finalization.
ExecuteRunOptionsRunAgentOptions & { idempotencyKey?; contextInit?; memory?; session?; rethrow?; sinks? }memory: at run start, hydrate store.load(keys) into the blackboard (user seed wins), and write back at the end; session: conversation-history read/write (orthogonal to the former).
Runclass: start() / finish(result) / fail(error) / toMeta()The run state machine (queued → running → succeeded/failed); runId == recorder.traceId.
RunContextclass: static current() / get(key) / set(key, value) / has / delete / keysThe in-run blackboard, propagated via AsyncLocalStorage, with no global singleton. Key types narrow through the Blackboard declaration merge (undeclared = string).
withRunContext(ctx: RunContext, fn) → T | Promise<T>Runs fn within the given ctx scope (the return type follows fn: a sync fn gives T).
SuspendedReason'approval' | 'timer'The suspension reason: "waiting for a human decision" (approval, HITL) and "waiting for an instant" (timer, durable timer). The decisions (whether approvalTimeoutMs expires, whether approve can wake it) rest on the reason rather than the status — so mishaps like "a time suspension woken early by an approval timeout" are unrepresentable in types.
RunMeta / RunStatus{ runId; status; idempotencyKey?; createdAt; startedAt?; finishedAt?; error?; suspendedReason?; wakeAt? } / 'queued' | 'running' | 'suspended' | 'succeeded' | 'failed' | 'cancelled'The run record and status. suspended is non-terminal: it does not occupy a concurrency slot, resumePending does not pick it up, and awaitTask keeps waiting; what it waits for is stated by suspendedReason ('approval' = waiting for a human decision, 'timer' = waiting for an instant), with wakeAt the target instant for the latter. cancelled (cancelled, the host called runner.cancel) is terminal and separate from failed: cancellation is not failure (the engine-side aborted finalization semantics), and operational readings should not lump "someone pressed it" together with "it fell over"; the mechanism is the same abort signal, and the difference is intent — a runTimeoutMs expiry still lands as failed.
SystemPromptclass: new SystemPrompt({ version? }); add(name, text, stable?) / add(section) / build({ cache? })Assembles system in sections: a stable section goes into the cacheable prefix and gets an ephemeral breakpoint, a volatile section goes after it; version is automatically written to the run root's system.version.
SystemPromptOptions / SystemSection{ version? } / { name; text; stable? }The constructor options and the section; stable=false does not enter the cacheable prefix.
SessionStore / InMemorySessionStoreinterface: load(id) / append(id, messages)Conversation history across runs (append-only), orthogonal to MemoryStore (a key-value blackboard); they can be used together; only turns that run successfully are written back.
normalizeMessages(input: RunInput | unknown) → MessageParam[]Normalizes an arbitrary task input (string / messages / { prompt | text | messages }) into messages; an empty array throws.
RunInputstring | MessageParam[] | { prompt?; text?; messages? }The raw input shapes the transport layer accepts.
RunSpec{ messages; options?; source? }The normalized input for one task; source marks the trigger origin (sync / async / schedule:<id>).
RunInvocationOptions{ model?; fallbacks?; traceContent?; labels?; maxTokens?; maxIterations?; client?; onText?; signal?; blackboard?; contextPolicy?; retry?; idempotencyKey?; traceLimits?; onTraceEvent?; traceContext?; rethrow?; tools?; maxTotalTokens?; maxCostUsd?; priceOverrides?; toolTimeoutMs?; maxToolConcurrency?; maxEventChars?; approvals?; events?; sessionId? }The invocation parameters for a single run, shared by all three trigger kinds (excluding the non-serializable store / session, which are only available for in-process direct calls). fallbacks is the model fallback chain (a ring's client is an in-process object; a ring that only writes model still holds across a restart-and-resume). events are run events delivered during a suspension (persisted with the TaskRecord; the engine injects them into the message stream on the resumed segment).
runSync / createSyncHandler(app, input, opts?) / (app) → (input, opts?) => …Synchronous RPC: normalize the input then call app.run directly.
AsyncRunnernew AsyncRunner(app, { client?, store?, concurrency?, maxQueued?, runTimeoutMs?, approvalTimeoutMs?, taskSinks? }): submit / approve(taskId, decisions, opts?) / signalTask(taskId, event) / cancel(taskId) / resumePending(opts?) / awaitTask(taskId, opts?) / drain(opts?) / inFlight / isDrainingSubmitting returns a task record and executes in the background; tasks with the same idempotency key that have not failed are deduplicated last-wins. maxQueued (default 0 = unbounded) is the cap on tasks waiting for a slot — over it, submit throws TaskQueueFullError and HTTP returns 503 + Retry-After (the same pattern as maxConcurrentRuns; the resume path is not gated, only new work). approve (HITL) approves a suspended task: idempotent per tool_use_id (the first decision wins), and once all decisions are in it persists first then resumes execution; it accepts only ids in that task's pendingApprovals (one extra → 400, the whole batch is rejected, the record is untouched). signalTask (event intake) delivers a TaskEvent to a suspended task: only effective for suspended, persisted first then dispatched, idempotent dedup on eventId (a repeat → 409); the pending-injection buffer is capped at 64 entries (when full → 409, this one is not accepted). approvalTimeoutMs (default 0 = unbounded) is a lazy approval timeout: when approve / poll / resumePending reads an expired suspended task it auto-rejects it ("approval timeout") and redispatches, with no timer started. drain() = reject new work → wait for in-flight work to finish (returns false on timeout, unfinished tasks stay in the store). awaitTask() waits for a given task to reach a terminal state: when this process writes the terminal state it is woken by an event, with intervalMs (default 250ms) only a fallback poll (needed only for an async store whose terminal state is written by another process).
AsyncRunnerOptions / ResumePendingOptions{ …; runTimeoutMs?; approvalTimeoutMs?; taskSinks?; onPersistError? } / { staleAfterMs? }runTimeoutMs truly aborts the in-flight run on expiry; approvalTimeoutMs is the lazy timeout for suspended approvals (no timer started); resumePending uses ownerId to skip this process's records (suspended is never picked up — a suspension is not an orphan), and staleAfterMs is only a fallback heuristic for when the owner's liveness cannot be determined: when it can, liveness is judged by lease — if the ownerId's hostname is this machine, it checks whether that pid is still alive (a live one is never stolen; a dead one is immediately stealable = a crashed orphan does not wait out a freshness window), while a foreign host / old format falls back to "does this record look old enough"; approvalTimeoutMs holds only for suspendedReason === 'approval' (a time suspension is not woken early by an approval timeout).
TaskSink{ onFinished(rec: TaskRecord) }A task-completion callback: once terminal, delivered one by one with await (a throw is swallowed), notified before the in-flight count is decremented.
PersistFailureInfo{ record: TaskRecord; error: unknown; phase: 'initial' | 'outcome' }The shape of a persist failure (the callback argument of AsyncRunnerOptions.onPersistError). This failure used to be swallowed by #safeSave with no outlet, so the host could not do what it was told to do. The costliest case is phase:'outcome': when the terminal-state write fails while the store sits at running though the run has actually finished (side effects already happened) — after a restart resumePending will treat it as an orphan and re-run it, so side-effecting tools must be idempotent themselves. Same cause and shape as createOtlpExporter({ onExportError }).
AppCallable{ name; run(messages, opts?) }The minimal application call surface.
Scheduler / ScheduleHandle / ScheduleEveryOptionsnew Scheduler(runner): every(intervalMs, input, opts?) / at(when, input, opts?)Periodic / one-off dispatch, submitted as an async task when due; returns a ScheduleHandle { id, cancel() }; an illegal interval / date throws outright. every's maxInFlight (default 1) is "the cap on tasks not yet terminal", and Infinity = gate open; 0 / a negative number throws at construction — the gate check size >= maxInFlight is always true at 0, so a periodic task would never dispatch and never report an error.
TaskStoreinterface: save / get / byIdempotency / list / clear (optional listDue(before) / compact() / close())The task-record storage surface (sync or async implementations both fine, returning MaybePromise); byIdempotency is the basis for last-wins dedup. listDue (a due index, optional): returns only records "suspended with wakeAt ≤ before"; AsyncRunner's due-wake scan uses it if present and falls back to a full scan otherwise (semantics unchanged); it is worth implementing for a store backed by real queries (sqlite), and unnecessary for one that keeps records in memory (memory / file).
TaskRecord{ taskId; status; idempotencyKey?; spec: RunSpec; runId?; ownerId?; approvals?; pendingApprovals?; pendingEvents?; deliveredEventIds?; createdAt; startedAt?; … }The task record; runId (== traceId) is backfilled on completion. ownerId is an internal identifier (p<pid>@<host>-<8 hex>, carrying a hostname since 2026-09-28) — resume claiming relies on it to decide whether the owner is still alive, and a host should not parse its format. HITL: when suspended, spec.messages holds the extended history (ending with an assistant message containing the unresolved tool_use), pendingApprovals is the list of pending tool_use_ids, and approvals are the persisted decisions (survive a restart). Event intake: pendingEvents are events received during the suspension and awaiting injection into history (cleared once the run goes through), and deliveredEventIds is the eventId idempotency bookkeeping (bounded FIFO).
InMemoryTaskStoreclass (TaskStore): { maxRecords? }A Map implementation for tests and default scenarios; on exceeding the cap it evicts starting from the oldest terminal record (in-flight records are never evicted).
FileTaskStorenew FileTaskStore(file)Persists one snapshot per line as JSONL; after a host restart resumePending() continues queued/running tasks.
SqliteTaskStorenew SqliteTaskStore(path)node:sqlite persistence (WAL, reads and writes not mutually exclusive), supporting ':memory:'. It implements the optional listDue(before) due index (a derived wake_at column + a (status, wake_at) index; an existing store is migrated in place and backfilled at construction) — the due-wake scan is O(due count) rather than O(full table).
RedisTaskStorenew RedisTaskStore(client, { prefix?, ttlSeconds? })A duck-typed Redis client (ioredis/node-redis both fine, an optional peer); TaskStore methods may return a promise.
RedisLike / RedisTaskStoreOptions / RedisSetOptions{ get; set; del; expire?; keys | scanIterator } etc.The minimal shape and option types for a Redis client (the framework does not import the redis package). TTL is applied via expire(key, seconds) (same name and shape across both clients); set takes only two arguments — the trailing options argument has opposite shapes between the two, and picking either one silently fails on the other.
MaybePromise<T>T | Promise<T>The unified shape for sync / async implementations (a caller's await works with both).

run · hosting and exports

The same application wires into HTTP in one line (with an auth seam, graceful shutdown and health checks); traces export to any OTLP collector or a custom sink; an OpenAI-compatible endpoint wires in multiple models in one line.

ExportSignature summaryNotes
createHttpHandler(app: AppCallable, opts?: HttpHandlerOptions) → HttpHandlernode:http integration: POST /run (sync / SSE), POST /tasks (async), GET /tasks/:id (poll), POST /tasks/:id/approve (HITL approval: 200 returns the task record; 404 if absent / 409 if the state is wrong / 400 for an illegal body, or when decisions contains an id not in this task's pendingApprovals — the whole batch is rejected, the record untouched; still available while draining), POST /tasks/:id/events (event intake: the allow-listed { eventId?, type, payload } delivered to a suspended task; same 404 / 409 / 400 / 413 tiers; still available while draining), GET /healthz. The return value can be passed straight to http.createServer, with drain() and runner also attached. The SSE side has backpressure protection: once the downstream backlog exceeds sseMaxBufferedBytes the stream is closed and the corresponding run aborted.
HttpHandlerOptions{ runner?; authenticate?; metrics?; maxConcurrentRuns?=32; maxBodyBytes?=1MiB; exposeErrors?=false; sseMaxBufferedBytes?=8MiB }authenticate intercepts at the entry, before the body is read (all paths except /healthz and /metrics); passing metricsSink() (or any { render() } / a closure returning a string) gives you a built-in GET /metrics (unauthenticated, same tier as /healthz); over concurrency returns 503 + Retry-After; with exposeErrors off a 500 returns only a generic message; sseMaxBufferedBytes is the SSE downstream backlog cap, and over it the stream is closed and the corresponding run aborted (guarding against a client that "connects but does not read" eating all memory). ⚠️ maxBodyBytes and sseMaxBufferedBytes must be positive (0 throws at construction): 0 would make every request with a body a 413 / every stream not survive a single frame — use Infinity for "unbounded".
HttpHandler((req, res) => Promise<void>) & { drain(opts?) → Promise<boolean>; runner }Shutdown: drain() rejects new work → waits for in-flight work → force-closes SSE streams; unfinished tasks stay in the store. A live connection (SSE) does not get an unlimited window: when timeoutMs is 0 (wait forever), close first then wait — otherwise that wait has no endpoint. The return value reports the two finalizations faithfully: if it cut off an in-flight run (a /run stream, finalized by abort) then it did not finish cleanly (false); only closing bystander long connections (/tasks/:id/stream) loses no work and does not affect the return value.
HttpExceptionnew HttpException(status, body)Throw it in an auth hook to respond with its status / body (throw 403 to return a 403); throwing anything else returns 401 with the raw text only in the server log.
HealthResponse{ ok; inFlight; uptimeMs; draining; suspended: { approval; timer; nextWakeAt } }The GET /healthz return: unauthenticated, and returns 200 even while draining; ok is always true (if it can respond, the process is alive), and readiness is read from draining. suspended is the suspension reading (counts grouped by reason + the earliest target instant; nextWakeAt is null with no time suspension) — the scope is the records this process can see, from the same table as inFlight.
RunHttpResponse / TaskSubmitBody{ runId; …result } / { input; options?; idempotencyKey? }HTTP contract types.
registerDefaultTraceSink(sink: TraceSink) → voidRegisters a global default sink (snapshotted and merged at createApp construction; already-built apps are unaffected by later registrations).
createOtlpExporter({ endpoint, headers?, serviceName? }) → { export(trace) → Promise<void> }Exports traces to any OTLP/HTTP collector (default service.name is 'agentia'); nested spans are flattened into attributes + events, zero dependencies. Attributes align with OTel GenAI semconv v1.37 (additive: appends gen_ai.* keys, keeping the old ones), and a score event translates to gen_ai.evaluation.result.
jsonlTraceSink(options: { path: string }) → TraceSinkA JSONL file sink: one bare Trace per line appended to disk — the artifact is exactly the input file for agentia report / agentia diff / agentia harvest. A missing parent directory is created recursively at construction; it appends rather than overwrites; a write failure is thrown to the caller (swallowed by flushSinks with a console.warn).
createOpenAIClient(opts?: { apiKey?; baseURL?; fetchImpl?; stream? }) → ModelClientAdapts an OpenAI-compatible endpoint (including DeepSeek and others) into a ModelClient: true streaming, image blocks converted to image_url, forwards signal. stream: false falls back to a one-shot request.
createAnthropicClient(opts?: { apiKey?; baseURL?; … }) → ModelClientThe default ModelClient (Anthropic Messages API). To customize, pass only apiKey / baseURL (defaults come from env vars) — no direct dependency on the vendor SDK.
AnthropicClientOptions{ apiKey?; baseURL?; … }The configurable options for createAnthropicClient (other parameters are passed through to the SDK verbatim).
MemoryStoreinterface: load(keys) / save(entries) + optional loadWithRev / saveIfRevThe cross-run memory storage surface (sync or async implementations). With only load/save it is last-write-wins: concurrent runs on the same keys lose writes and the framework cannot detect it; to be able to notice, implement loadWithRev + saveIfRev as a pair (CAS: on conflict the framework does not write and speaks up, it does not merge for you). Implementing only one is an assembly error → the entry throws a TypeError.
InMemoryMemoryStoreclass (MemoryStore)A Map implementation: in-process, no persistence. It is versioned (one monotonic counter for the whole store, coarse-grained → it over-reports conflicts rather than missing them).
MemorySnapshotinterface: values / revThe return of loadWithRev: values as in load; rev is an opaque version handle (the framework does not interpret it, only passing it back verbatim to saveIfRev when this run writes back). The granularity is up to the store, the finer the fewer false conflicts.
MemoryWriteResultinterface: committed / reason?The result of saveIfRev: only committed: true means it was actually written; on false (reason: 'conflict') not a single character may be written, and the framework logs a warning accordingly ("write-back rejected" is no longer silent).

integrations · MCP and metrics

The external-system adapter layer, with zero runtime dependencies — the MCP bridge and optional capabilities (zod / Redis) are all duck-typed (the same paradigm as RedisLike); the MCP connectors ship in the box yet use only the standard library, see below.

ExportSignature summaryNotes
mcpTools(client: McpClientLike, opts?: McpToolsOptions) → Promise<AgentTool[]>Maps an MCP server's tools/list into framework tools that go straight into createApp({ tools }); they share the pool with local @Tool (same middleware, same duplicate-name check).
McpClientLike{ listTools(); callTool(name, args) }The minimal shape for an MCP client: any object implementing it can be plugged in, and the framework does not import the MCP SDK; use this seam for the official SDK / a remote server / a custom transport. Have callTool throw on failure (the framework wraps it as is_error).
McpConnectorMcpClientLike & { close() }The common surface of the two shipped connectors. close() is idempotent; on the stdio side returning means the child process has terminated (SIGTERM → grace period → SIGKILL, still waiting for 'exit', leaving no orphan); HTTP does a best-effort session DELETE. Calling after close throws a readable error.
createStdioMcpConnector(cmd: string[], opts?: StdioMcpConnectorOptions) → McpConnectorThe shipped stdio connector: spawns a child process speaking newline-delimited JSON-RPC (initialize → notifications/initialized → tools/list / tools/call). Lazy (no spawn at construction); it catches three things only a connector can: the async 'error' event of a failed spawn, stdout framing, and turning a protocol-level isError: true into a thrown error. Uses only node:child_process.
createStreamableHttpMcpConnector(url: string, opts?: StreamableHttpMcpConnectorOptions) → McpConnectorThe shipped StreamableHTTP connector: one endpoint taking JSON-RPC POSTs, accepting both application/json and text/event-stream responses; it remembers Mcp-Session-Id, sends it back on subsequent requests, and best-effort terminates the session on close(). Session-expiry self-healing: receiving a 404 with a session id ⇒ discard the session → re-handshake → retry this call once (only once; onSessionExpired is observable). An HTTP failure with a numeric status is automatically classified as retryable / non-retryable. Uses only the global fetch.
StdioMcpConnectorOptions{ env?; cwd?; stderr?; clientInfo?; protocolVersion?; timeoutMs?=60000 }timeoutMs applies only during assembly (handshake + tools/list) — those two steps have no other referee, and a stuck server would leave createApp hanging forever.
StreamableHttpMcpConnectorOptions{ headers?; clientInfo?; protocolVersion?; timeoutMs?=60000; onSessionExpired?; fetchImpl? }headers is where auth goes; onSessionExpired is called once on session-expiry self-healing (the healing itself is silent — use this to count / alert); fetchImpl is the test injection point (same pattern as createOpenAIClient).
McpToolInfo{ name; description?; inputSchema? }A shape-subset of a tools/list entry; inputSchema is already JSON Schema → used as input_schema as-is.
McpToolsOptions{ prefix?; server?; timeoutMs?=60000 }The default prefix is mcp_<server>_ (mcp_ when no server is given); timeoutMs is the fallback single-call timeout (give up waiting, not cancel) — when the engine sets toolTimeoutMs, this option does not take part in the decision.
MCP_DEFAULT_TIMEOUT_MS60000The bridge's fallback single-call timeout (does not take part in the decision when the engine sets toolTimeoutMs).
MCP_CLOSE_GRACE_MS2000The grace period in close() between SIGTERM and SIGKILL.
createMcpServer(app: McpServerApp, opts: McpServerOptions) → McpServerThe MCP reverse bridge (R8-P5): exposes an app's capability menu (the assembled one, post-middleware) as an MCP server — Claude Code / Cursor and any MCP host can call it directly. Transport is stdio (newline-delimited JSON-RPC) or http (StreamableHTTP: POST takes the payload and replies application/json; initialize mints an mcp-session-id header but does not validate it; GET → 405, DELETE → 200; a client disconnect aborts that call's signal). The scope is tools only (initialize / tools/list / tools/call + ping). Every tools/call builds one trace (run root mcp.tools/call + a capability span + tool.input/tool.output events) delivered to opts.sinks. Standard library only.
McpServerApp{ tools: AgentTool[] }The app input for the reverse bridge (duck-typed): an AgentApp satisfies it — app.tools is the assembled menu.
McpServerOptions{ transport: 'stdio' | 'http'; host?; port?; path?='/mcp'; server?; client?; sinks?; toolTimeoutMs?; auth?; name? }client defaults lazily to createAnthropicClient() (constructed only on the first tools/call); server mounts onto an existing http.Server (it does not listen for it; close only detaches the handler); auth is the http-side auth hook (before the body is read, throwing gives 401 — same discipline as the HTTP host; stdio trusts the parent process); toolTimeoutMs has the same semantics as the engine (timeout = give up waiting, returning isError).
McpServer{ url: string | undefined; ready: Promise<void>; close(): Promise<void> }The reverse-bridge handle: url is the actual endpoint in http mode (always undefined for stdio); close() is idempotent and aborts in-flight calls' signal; it only closes a server it owns, and for a mounted one it only detaches the handler.
metricsSink(opts?: MetricsSinkOptions) → MetricsSinkAn in-process metrics accumulator that naturally satisfies TraceSink → createApp({ sinks: [metricsSink()] }) wires it in with no new outlet. Every dimension is derived from the existing trace (no instrumentation needed): run level; capability level (per tool/skill/subagent: call count, error count, latency, tokens, cost — tools read the tool.output event on the turn, skill/subagent read the capability span); model level (attributed by model, plus model_unpriced_turns_total and model_usage_missing_turns_total — the latter is "the upstream returned no usage", so cost will look like 0, counted separately from "the model has no price"); score level (agentia_score gauge + agentia_score_total counter, from the score event on the run root); and agentia_dropped_keys{kind="capability"|"model"|"score"} — the count of distinct keys folded by the cardinality cap, always emitting three samples (folding is not silent); with labelKeys enabled it adds kind="label:<key>" per key, and kind="label:combos" when the combination count hits its cap (when this is off those rows are entirely absent — not 0, but rather that dimension was never enabled).
MetricsSinkTraceSink & { snapshot(); render(); contentType?; flush(); stop(); reset() }snapshot gives numbers (including capabilities / models / droppedCapabilities / droppedModels / droppedScores / exemplars), render gives text (/metrics returns it directly), contentType declares the format of the render product (follows the export mode; the built-in route reads it to set the response header), flush exports once on demand (OTLP mode), stop halts scheduled export, reset clears the accumulators. The count of distinct folded keys is visible on both sides: snapshot() (three dropped* fields) and render() (agentia_dropped_keys).
MetricsSinkOptions{ export?='prometheus' | 'openmetrics' | 'otlp'; endpoint?; intervalMs?=60000; resourceAttributes?; serviceName?; timeoutMs?; onExportError?; windowSize?=1024; prefix?='agentia_'; labelMode?='capability' | 'kind' | 'none'; maxCapabilities?=200; maxModels?=50; maxScores?=200; labelKeys?=[]; maxLabelValues?=100; maxLabelCombos?=200; buckets? }The token measure = the sum of all four kinds (consistent with BudgetGuard). Durations are given in two measures side by side: a histogram (*_bucket/_sum/_count, cumulative semantics, aggregatable across instances) and an in-window exact quantile (a gauge, easy to read on a single instance). export:'openmetrics' emits OpenMetrics text with exemplars (metric ↔ trace jumps); export:'otlp' hand-writes OTLP/JSON with zero dependencies and must be given an endpoint (a missing one throws at construction). labelMode + maxCapabilities/maxModels/maxScores cap those three dimensions (model and score keys can be unbounded just like capability keys) — over the cap, keys are folded into __other__ (the label name per outlet is capability / model / name respectively); buckets still accumulate so totals are not lost, only label granularity is. labelKeys (R8-P4 attribution labels) explicitly names which keys in the run root's labels.* go onto metric labels (none by default); each key's distinct-value count is capped by maxLabelValues (overflow folds into __other__). maxLabelValues caps each key's value domain, but what enters memory is the combination of keys (a cross product: a domain of 100 with 3 keys is a million entries) — maxLabelCombos (default 200) caps the combination count, folding over-limit new combinations into a single all-__other__ bucket (volume still collected, only label granularity lost), with the folded count visible in snapshot().droppedLabelCombos and agentia_dropped_keys{kind="label:combos"}; this cap is likewise not disableable (the memory invariant requires every new cardinality dimension to have a cap; opt-in cannot stop "we know there are thousands of tenants but we insist on shipping it"). Under a single-key configuration the combination count ≈ the value count, so behavior is unchanged.
MetricsSnapshot / CapabilityMetrics / ModelMetrics / RunLabelMetrics / ExemplarSnapshot{ runs; failed; latencyP50; latencyP95; tokens; costUsd; capabilities; models; scores; runLabels; droppedCapabilities; droppedModels; droppedScores; droppedLabelValues; droppedLabelCombos; exemplars }capabilities gives calls/errors/quantiles/tokens/costUsd per capability label (a tool has no token semantics, so it is null); models gives turns/tokens/costUsd/unpricedTurns/usageMissingTurns/quantiles per model; scores gives the latest value / count / total per scoring dimension (key name@source, from the score event of attachScore); runLabels (R8-P4) gives runs/failed/tokens/costUsd per attribution-label combo (key = k=v comma-joined) — only for keys named by labelKeys. dropped* is the count of distinct keys folded into __other__ (each remembers at most 1024; beyond that it is a lower bound; droppedLabelValues is counted per labelKey and droppedLabelCombos is the batch folded by the combination cap). exemplars (ExemplarSnapshot) is the metric ↔ trace jump: failed = the most recent failed run, slowest = the traceId/spanId of the slowest run so far (attached on the OpenMetrics and OTLP outlets).
DEFAULT_BUCKETSreadonly number[]The default bucket boundaries for the duration histogram (ms): 25 / 50 / 100 / 250 / 500 / 1k / 2.5k / 5k / 10k / 30k. Overridable via buckets (must be strictly ascending).
buildRunReport(trace: Trace) → RunReportA tuning report (a pure function): capabilities ranked by total duration descending (calls / errors / total / max / tokens / cost), model attribution, and the list of unpriced models — answering "which knob should I turn". A single run is a small sample, so total / max lead and no quantiles are given.
mergeRunReports(reports: RunReport[]) → RunReportCross-run aggregation (quantiles only carry statistical meaning across multiple runs): calls/errors/durations accumulate, capabilities are re-ranked by the new total duration, and unpriced models are unioned.
renderRunReport(report: RunReport) → stringA human-readable plain-text table (for the CLI / logs); unpriced models are explicitly marked in the cost column. agentia report <trace.jsonl> is its thin shell.
RunReport / CapabilityReport{ traceId; status; durationMs; totalUsage; models; capabilities; unpricedModels; usageMissingModels; runs }The report structure; CapabilityReport carries durations (the raw samples, on which merge recomputes cross-run quantiles).
ModelReport / DurationReport{ model; turns; tokens; tokensTotal; costUsd; unpricedTurns; usageMissingTurns; durationMs; durations } / { total; max; p50; p95 }Model attribution and duration aggregation (a single run's p50/p95 is for reference only).

eval · regression assertions

Makes scripted model round trips a first-class capability, so "did anything regress after changing a prompt / model / adding a tool" becomes an assertion rather than an eyeball check. The assertion source is the existing Trace.

ExportSignature summaryNotes
scriptedClient(steps: ScriptedStep[]) → ModelClientReturns model responses in script order and really emits the text blocks through on('text') (the onText / SSE path runs the real route). A step advances only after finalMessage succeeds → a throwing step replays the same step on retry.
ScriptedStepRecord<string, unknown> | (params) => messageOne scripted step: give the response object directly, or compute it from this request's parameters (may throw to simulate a 429 / network failure). Exhausting the script throws.
defineEval<T>(def: EvalDefinition<T>) → { name; run() → Promise<EvalReport> }Defines "cases + assertions". run() does not throw (a failed case goes into the report, so one pass shows all regressions); only a failure to build the app bubbles up.
EvalDefinition{ name; app; cases; expect }The app is built once per run (assembly reused across cases); expect throwing = that case failed.
EvalCase{ name?; input; client; opts? }One case; opts is passed through to app.run (inject resultSchema to assert on result.typed).
EvalContext{ trace: Trace }The assertion context (the same one as result.trace).
EvalReport{ name; total; passed; failed; ok; cases: EvalCaseReport[] }The report; ok is true only if all pass (total=0 counts as passing). For CI, just read report.ok.
EvalCaseReport{ name; ok; stopReason?; error?; trace? }A failed case carries the full trace — that is what you look at to diagnose a regression.

container · explicit DI

Explicit provider registration with override semantics (later registrations override earlier ones); resolve caches singletons and detects circular dependencies; re-registering transitively invalidates downstream caches.

ExportSignature summaryNotes
Containerclass: register(...providers) / has(token) / resolve<T>(token)The DI container; register returns this for chaining.
TokenstringThe provider registration key.
ProviderValueProvider | ClassProvider | FactoryProviderThe union of the three provider shapes.
ValueProvider{ provide; useValue }Provides a value directly.
ClassProvider{ provide; useClass; deps? }Class instantiation; deps are injected as constructor parameters in order.
FactoryProvider{ provide; useFactory; deps? }A factory function; deps correspond one-to-one with useFactory's parameters (parameter types are free and will not be rejected for contravariance).

toolkit · declarative capabilities and assembly

Decorators declare the four capability kinds, and createApp assembles: DI registration → scan and collect the menu → static validation at assembly time (menu duplicate check, tools-reference existence, toolSources targets, DI circular dependencies).

ExportSignature summaryNotes
createApp(opts: AppOptions) → AgentApp; with discover → Promise<AgentApp>Assembles the app; directory discovery uses dynamic import, hence the Promise return.
AppOptions{ name?; providers?; modules?; discover?; system; model?; fallbacks?; traceContent?; labels?; maxTokens?; maxIterations?; contextPolicy?; retry?; maxTotalTokens?; maxCostUsd?; priceOverrides?; onUnpricedModel?; toolTimeoutMs?; maxToolConcurrency?; maxEventChars?; toolSources?; tools?; middleware?; sinks?; onTraceEvent?; traceLimits? }system is required: a SystemPrompt instance (auto-cached) or an assembled SystemParam. tools are bare tools going straight into the main menu (dynamic tools like MCP bridges go here, sharing middleware and the duplicate check); sinks is the trace outlet. priceOverrides / onUnpricedModel are the default cost configuration (overridable per run).
AgentAppclass: run(messages, opts?: RunAppOptions) → Promise<AgentRunOutput>; tools; containerThe assembly product; each run rebuilds system from SystemPrompt to keep volatile content fresh.
RunAppOptionsRunInvocationOptions & { system?; resultSchema?; session? }Per-call parameters: override system, give a structured-result schema, pass a session store (session is not serializable, so it is available only for in-process direct calls).
AgentRunOutput{ run: Run; result: AgentRunResult }The return of app.run.
defineModule / AgentModule(m: AgentModule) → AgentModule / { providers; middleware? }The third-party capability-package convention: providers register before the app level (an app-level registration on the same token overrides), middleware is composed on the outer layer; defineModule is an identity function.
Tool@Tool(spec: ToolSpec) method decoratorA class method is a tool: parameters are validated against the schema before execution.
ToolSpec{ name?; description; schema: JsonSchema; strict?; approval?: 'required' }name defaults to the method name; strict requires additionalProperties:false + a complete required list; when schema is fromZod<T> the method signature is validated; declaring approval (HITL) makes each call suspend and wait for human approval.
collectTools(instance: object) → AgentTool[]Scans @Tool methods along the prototype chain and binds them to the instance.
Skill@Skill(spec: SkillSpec) method decoratorThe method body is a deterministic script; a model call happens only via an explicit ctx.llm() and is recorded under the skill's own capability span.
SkillSpec{ name?; description; schema?; model?; maxTokens?; maxIterations?; tools? }tools is a list of container provider tokens (callable in ctx.llm()).
SkillContext{ model?; llm(opts: SkillLlmOptions) → Promise<SkillLlmResult> }A restricted sub-run handle: each llm() opens an independent agent loop under the capability span.
SkillLlmOptions / SkillLlmResult{ prompt?; messages?; system?; model?; maxTokens?; maxIterations? } / { text; stopReason }prompt and messages are mutually exclusive; system accepts a string or a SystemPrompt.
collectSkills / skillToTool(instance) → SkillCapability[] / (capability, resolveTools) → AgentToolScans @Skill methods / compiles them into main-agent menu entries.
SkillCapability{ name; description; inputSchema; spec; invoke(input, ctx) }Carries the bound-instance method executor.
SubAgent@SubAgent(spec: SubAgentSpec) method decoratorThe method body does not execute: the framework starts a separate isolated loop from system, intermediate steps do not leak out, and only the final report flows back.
SubAgentSpec{ name?; description; schema; system; tools?; model?; maxTokens?; maxIterations?; resultSchema? }system accepts a string / SystemPrompt / (task) => SystemParam (assembled dynamically, with nothing appended by the framework); tools is a list of provider tokens. A reference cycle throws at assembly time (a direct self-reference, or one that loops back through another provider, both count): reference resolution is lazy, so a cycle only exists at runtime, while maxIterations caps only one turn's width — depth has no gate at all ⇒ a single self-call is unbounded recursion. The error message gives the node sequence on the cycle.
collectSubAgents / subagentToTool(instance) → SubAgentCapability[] / (capability, resolveTools) → AgentToolScans @SubAgent methods / compiles them into an AgentTool (the sub-loop reuses ctx.client and recorder).
SubAgentCapability{ name; description; inputSchema; spec }Sub-agent capability metadata.
Prompt@Prompt(spec: PromptSpec) method decoratorA plain-text asset compiled into a side-effect-free pull-style tool; recomputed on each call (volatile semantics).
PromptSpec{ name?; description; schema?; version? }Write description so it is clear when to pull, and the model decides based on it; schema defaults to an empty object (a parameterless asset); version is the asset version — collected by capability name at assembly time and recorded on the run root's prompts.versions attribute on each run.
collectPrompts(instance: object) → AgentTool[]Scans @Prompt methods (static methods are fine too).
asset(base: string | URL, rel: string) → stringasset(import.meta.url, './system.md') reads a text asset relative to the capability directory; it reads fresh each call and does not cache.
loadEnvFile(options?: LoadEnvOptions) → Record<string, string>Reads a .env into process.env and returns the keys that actually took effect this time. The framework does not read .env automatically — which file and when is decided by the host's startup code (the scaffold's main.ts has it built into the first line). A missing file silently returns {}; real environment variables win, and defined keys are not overwritten (including the empty string).
LoadEnvOptions{ path?; override? }path defaults to .env (relative to cwd); override defaults to false (set it true to let the file beat environment variables).
discoverProviders(dir: string | string[]) → Promise<Provider[]>Give one directory or a set of directories (array order is assembly order) and it scans each for <name>/index.ts: the default export may be a class / Provider / Provider[], a class registers under the folder name as token; if any directory in the array does not exist it errors, and a duplicate token across directories leaves a warning (the later one overrides at assembly time).
applyMiddleware(tools: AgentTool[], middleware: CapabilityMiddleware[]) → AgentTool[]Wraps the whole menu in an onion model (chain order = registration order, first registered is outermost); an empty chain returns as-is at zero cost.
CapabilityMiddleware(call: CapabilityCall, next: CapabilityNext) => unknownCross-cutting around a capability call (auth / rate limiting / caching / audit / quota); not calling next short-circuits, and a throw is handled as a capability failure.
CapabilityCall / CapabilityNext{ capability; input; ctx? } / (input?) => unknownThe input schema has already been validated; next(newInput) may rewrite the input; calling it twice in a row throws.
fromZod(jsonSchema: JsonSchema, zod: unknown) → JsonSchemaOptional zod integration (a peer): validation goes through zod safeParse (shape-based detection; the framework does not import zod), and error paths are returned to the model as-is so it can self-correct.