DOCS · @migor/agentia
Guide
From a three-minute smoke test to production hardening: the four capability kinds, directory
assembly, middleware, trigger modes, long context, observability and hard cost
control. Every API and command matches docs/usage-guide.md word for word; for a
quick scan of the exported surface see the public API reference.
Quickstart
Agentia is a declarative framework for agent services: decorators declare four kinds of capability, the main agent acts as a router and orchestrates them, and one run hands back structured output plus a complete call tree (runId == traceId) with tokens / cost / metrics accounted for all the way down — replayable and auditable. Running in three minutes:
npm i -g @migor/cli
agentia create my-app # scaffold: four category folders + src/registry.ts + .env
cd my-app && npm install
agentia g subagent doc-reviewer # generate a capability and register it automatically
# put ANTHROPIC_API_KEY into the .env the scaffold generated (.gitignore already covers it)
npm run dev -- "review docs/report.md for me"
The entry point src/main.ts does exactly one thing: assemble the app, then trigger a run.
import { createApp, SystemPrompt } from '@migor/agentia';
const app = await createApp({
name: 'my-app',
discover: ['src/tools', 'src/skills', 'src/prompts', 'src/subagents'], // four category folders; order is assembly order
system: new SystemPrompt().add('role', 'You are the main agent of my-app; dispatch the capabilities on your menu as the task requires.', true),
});
const { result } = await app.run([{ role: 'user', content: process.argv[2] }]);
console.log(result.finalText);
Scenario guide
Find the recipe by what you want to do: five cards, each with a minimal skeleton and a link to go deeper. Every API and command matches docs/usage-guide.md word for word.
Ship an HTTP agent service
One handler serves synchronous POST /run (SSE when Accept: text/event-stream is set), asynchronous POST /tasks (de-duplicated by idempotency key) and the GET /healthz probe. Auth goes through the authenticate seam (it intercepts before the body is read) and shutdown goes through drain — unfinished tasks stay in the store, and resumePending() picks them up after a restart.
import { createServer } from 'node:http';
import { createApp, createHttpHandler, AsyncRunner, SqliteTaskStore, HttpException } from '@migor/agentia';
const app = await createApp({ name: 'my-app', providers, system });
const runner = new AsyncRunner(app, { store: new SqliteTaskStore('./tasks.db') });
await runner.resumePending(); // restart resumes queued/running, it does not discard them
const handler = createHttpHandler(app, {
runner,
authenticate: async (req) => {
if (!verify(req.headers.authorization)) throw new HttpException(401, { error: 'invalid credentials' });
},
});
createServer(handler).listen(8080);
// POST /run (sync / SSE) POST /tasks (async) GET /tasks/:id GET /healthz
process.on('SIGTERM', () => handler.drain({ timeoutMs: 15_000 })); // signal handling belongs to the host
See “Trigger modes” and “Host hardening”; for a complete runnable example (four capabilities + three trigger modes + auth + the full observability stack + a Dockerfile) see examples/complete.
Want to add a gRPC entry point as well: that is a host-side concern, and the framework does not build it in (there is exactly one rule — is the client in the standard library: the MCP connector can be built with spawn + fetch and adds no dependency, whereas gRPC would pull in @grpc/grpc-js). The recipe is in docs/usage-guide.md §6.2, and a host + client + end-to-end you can copy is in examples/grpc-host — it spells out the four semantics most easily missed: deadline / cancellation → signal, metadata traceparent → run-root link, framework errors → gRPC status codes, trace → sink.
Put monitoring on a run
Metrics and trace go through the same seam: metricsSink() satisfies TraceSink, and wiring it into sinks gives you the agentia_* metric family (four levels: run / capability / model / score, where scores come from attachScore as agentia_score / agentia_score_total); createOtlpExporter aligns with the OTel GenAI semconv on export (gen_ai.* keys, recognised directly by Langfuse / Grafana / Datadog).
import { metricsSink, createOtlpExporter } from '@migor/agentia';
const metrics = metricsSink();
const app = await createApp({
name: 'my-app', providers, system,
sinks: [metrics, createOtlpExporter({ endpoint: 'http://localhost:4318' })],
});
createHttpHandler(app, { metrics }); // built-in GET /metrics (unauthenticated, still answers while draining)
The Grafana dashboard JSON ships with the repo (examples/observability/grafana-dashboard.json, built against the default prefix of metricsSink — import and use). See “Observability” and “Ecosystem & observability”.
Offline evals and harvesting from production
defineEval + scriptedClient run offline in CI: scripted model round-trips with assertions written against the trace; each case automatically attachScores its verdict onto its trace (value 0|1, source is the eval name), which a downstream metricsSink can aggregate into a pass rate. For production incidents, agentia harvest turns them back into eval case skeletons.
import { defineEval, scriptedClient } from '@migor/agentia';
const report = await defineEval({
name: 'doc-review',
app: () => app,
cases: [{ input: 'Summarise this document', client: scriptedClient([searchMsg, submitMsg]) }],
expect: (r, { trace }) => { /* assert stopReason / tool sequence / result.typed */ },
}).run();
if (!report.ok) process.exit(1); // failing cases carry the full trace; in CI you only read report.ok
agentia harvest trace.jsonl --failed --limit 5 --out evals/harvested.ts # production trace → eval case skeleton
A skeleton is not a finished test: the trace does not record assistant text (so the script holds placeholders), and the pre-filled expect transcribes what actually happened (which is not the same as what should happen) — review it by hand before it enters CI. See “Ecosystem & observability” and “CLI reference”.
prompt / model A/B
Run the same input through two runs (swap the model or the SystemPrompt version) and let diffTraces surface the structural and cost differences. Without writing code, agentia diff compares two on-disk traces directly and exits 1 when they differ, so CI can block “the trajectory drifted after a prompt / model change”. To “re-run from turn N with a different question”, use forkMessages — it produces a set of messages you feed back into app.run, which starts a new run, not a continuation.
import { diffTraces, forkMessages } from '@migor/agentia';
const a = await app.run(input, { model: 'claude-sonnet-5' });
const b = await app.run(input, { model: 'claude-opus-5' });
const diff = diffTraces(a.result.trace, b.result.trace); // run-level summary + per-span field diff
// Forked replay: truncate before turn 3 of the main loop (0-based) and continue with a different question
const messages = forkMessages(a.result.trace, {
atTurn: 3,
append: [{ role: 'user', content: 'Different angle: give me the conclusion first, then the evidence.' }],
});
const c = await app.run(messages); // a new run, continuing from the real tool history before the fork point
agentia diff a.jsonl b.jsonl # compare two on-disk traces (same input shape as agentia report); exits 1 when they differ
See “CLI reference”; for the lossy boundaries (the trace records neither assistant text nor the original input) see the Known boundaries section and the A/B part of /llms-full.txt.
Human gate (HITL)
Middleware can await a human decision before letting a call through, and the engine will wait for the tool result: allow = next(); reject = do not call next() (short-circuit, the side effect never happens); reject and redirect the model = throw (the result is recorded as is_error and the run continues).
import type { CapabilityMiddleware } from '@migor/agentia';
const requireApproval: CapabilityMiddleware = async (call, next) => {
if (!DANGEROUS.has(call.capability.name)) return next();
const ok = await askHuman(call.capability.name, call.input); // the run simply waits here, for a person
if (!ok) throw new Error(`not approved: ${call.capability.name}`); // → is_error, the model takes another route
return next();
};
createApp({ system, providers, middleware: [requireApproval] });
Cross-process suspend / resume is supported: once you declare @Tool({ approval: 'required' }), every model call suspends the task as suspended (message history and the pending list are persisted with the task, so a restart loses nothing), and POST /tasks/:id/approve resumes from the breakpoint with that decision — see “Cross-cutting recipes” and the API reference.
Four capability kinds
Capabilities are what a service is made of. All four kinds sit on the main agent's menu as callable entries, share one namespace, and are de-duplicated at assembly time.
@Tool — function calls
A class method is a tool. Arguments are parsed against a JSON Schema and validated before execution; a failure comes back as is_error and the run keeps going.
import { Tool } from '@migor/agentia';
export default class Weather {
@Tool({
description: 'Get the weather for a city',
schema: {
type: 'object',
properties: { city: { type: 'string' } },
required: ['city'],
additionalProperties: false,
},
strict: true,
})
get_weather(input: { city: string }) {
return `Shanghai 72°F sunny (${input.city})`;
}
}
@Skill — code-driven flow
The method body is a deterministic script: whether to call the model, how many times, and how to post-process the result are all decided by code. Model calls happen only at an explicit ctx.llm(), accounted for as an llm.turn under the skill's own capability span.
import { Skill } from '@migor/agentia';
import type { SkillContext } from '@migor/agentia';
export default class WeeklyReport {
@Skill({
description: 'Ask the LLM for key points on a given topic',
schema: {
type: 'object',
properties: { topic: { type: 'string' } },
required: ['topic'],
additionalProperties: false,
},
})
async weekly_report(input: { topic: string }, ctx: SkillContext) {
const r = await ctx.llm({ prompt: `Give three key points on "${input.topic}"` });
return r.text; // the return value is the product, handed back to the main agent as a tool_result
}
}
@SubAgent — isolated sub-agents
Its own loop with a trimmed context. The method body never executes — the framework starts another agent from the system you gave it; intermediate work stays inside and only the final report flows back to the main context.
import { SubAgent, asset } from '@migor/agentia';
export default class DocReviewer {
@SubAgent({
description: 'Review a document independently, in the role set by system.md',
schema: {
type: 'object',
properties: { task: { type: 'string' } },
required: ['task'],
additionalProperties: false,
},
system: asset(import.meta.url, './system.md'),
})
// the method body never runs: only the method name and decorator metadata are read
doc_reviewer(_input: { task: string }): void {}
}
@Prompt — plain-text assets
Templates and playbooks, compiled into a side-effect-free fetch tool on the menu: the model calls it when it decides it needs the text, and the text is injected into context as a tool_result. Long text lives in a sibling .md file.
import { Prompt, asset } from '@migor/agentia';
export default class StyleGuide {
@Prompt({ description: 'Writing conventions (the description says when to fetch it)' })
style_guide(): string {
return asset(import.meta.url, './asset.md');
}
}
Layout & assembly
Four category folders, one folder per capability, long text in .md files; each capability is a class with a default export whose decorator declares what it is.
my-app/
├─ src/
│ ├─ main.ts # createApp assembly entry point
│ ├─ registry.ts # explicit registry (kept up to date by agentia g)
│ ├─ tools/weather/index.ts # @Tool
│ ├─ skills/weekly-report/ # @Skill
│ ├─ prompts/style-guide/ # @Prompt (with asset.md)
│ └─ subagents/doc-reviewer/ # @SubAgent (with system.md)
└─ tsconfig.json # include: ['src'] — one src covers everything
Two assembly routes
- Directory scan:
createApp({ discover: ['src/tools', …] })takes one directory or a list of them (order is assembly order), scans<name>/index.tsin each at startup, and registers a default-exported class under the folder name as its DI token (because it uses dynamic import, acreateAppwithdiscoverreturns a Promise). - Explicit registry:
createApp({ providers, system }), whereproviderscomes fromsrc/registry.ts(maintained automatically byagentia g).
Use either route or mix the two; for the same token the later one wins and the menu is never collected twice. Assembly runs static checks: duplicate menu entries, reference existence, DI cycles — all failing at startup rather than in production.
Middleware (CapabilityMiddleware)
Hooks around every capability call (tool / skill / subagent / prompt): auth, rate limiting, result caching, audit logging, timeout wrapping and other cross-cutting concerns live here, not inside business capabilities.
type CapabilityMiddleware = (call: CapabilityCall, next: CapabilityNext) => unknown;
interface CapabilityCall {
readonly capability: AgentTool; // name / description / inputSchema are readable
readonly input: unknown; // arguments that already passed schema validation
readonly ctx?: ToolRunContext;
}
// Pass through to the next layer; next(newInput) can rewrite the arguments, default keeps the current input
type CapabilityNext = (input?: unknown) => unknown;
The onion model
- Chain order = registration order: the first registered is the outermost layer, wrapping like an onion;
next()passes through,next(newInput)rewrites the arguments; not calling next short-circuits (result caching, for instance);- Throwing is treated as a capability failure: the engine wraps it into
is_errorfor the model and the run continues; - Wrapping happens at assembly time (in the AgentApp constructor) with zero intrusion into the engine; an empty chain passes straight through at no cost.
const app = await createApp({
name: 'my-app',
discover: ['src/tools', 'src/skills', 'src/prompts', 'src/subagents'],
middleware: [auth, cached], // auth is the outermost layer
system,
});
Example: auth
import type { CapabilityMiddleware } from '@migor/agentia';
// Sensitive capabilities require an authorisation marker in the execution context, otherwise fail as a capability error
export const auth: CapabilityMiddleware = (call, next) => {
if (call.capability.name.startsWith('admin:') && !call.ctx) {
throw new Error('forbidden: missing execution context');
}
return next();
};
Example: result caching
import type { CapabilityMiddleware } from '@migor/agentia';
const store = new Map<string, unknown>();
// A hit short-circuits (next is not called); a miss passes through and writes back
export const cached: CapabilityMiddleware = async (call, next) => {
const key = `${call.capability.name}:${JSON.stringify(call.input)}`;
if (store.has(key)) return store.get(key);
const out = await next();
store.set(key, out);
return out;
};
Trigger modes
Synchronous RPC, asynchronous tasks and scheduled dispatch share one input contract; swapping the host does not change the semantics.
Synchronous RPC
await app.run(messages) hands back the result and the trace directly — a good fit for request-response hosts.
Asynchronous tasks
AsyncRunner returns a task record as soon as you submit and runs it in the background; FileTaskStore snapshots one JSONL line at a time, and after a host restart resumePending() picks queued/running tasks back up.
const task = await runner.submit(
[{ role: 'user', content: 'Generate last week\u2019s ops report' }],
{ idempotencyKey: 'weekly-report:2026-W36' },
);
Scheduled dispatch
The Scheduler supports every (interval) and at (fixed time) schedules, turning a due schedule into an asynchronous task submission.
Idempotency
The async host has at-least-once semantics: submitting the same idempotencyKey again returns the existing record as long as the previous task has not failed (last-wins de-duplication); a failed task under the same key may be resubmitted as a new task.
Long-context policy
createBudgetPolicy provides a budget-driven context guardrail whose rules carry hysteresis, so it does not compress over and over on every turn:
- Estimate tokens for the whole message set; at or under budget → pass through unchanged (the fast path);
- Over budget, first do context editing: drop the oldest tool pairs, without extra model calls;
- Still over budget and a
summarizeis available → compaction: turn the old prefix into a summary and keep only the most recentkeepRecentitems; if fewer thancompactEveryturns have passed since the last compaction, skip (anti-flapping); - The framework does not invent tokens for you: the default token estimate is a CJK-aware heuristic, and you can inject one backed by
/count_tokensin production. - Only new messages are estimated: history is append-only, so the prefix estimate is reused, which makes each turn's estimation cost proportional to what this turn added rather than to the whole history.
import { createBudgetPolicy } from '@migor/agentia';
const policy = createBudgetPolicy({
budgetTokens: 60_000, // budget (estimated input tokens)
keepRecent: 20, // how many recent messages compaction keeps
editBeforeCompact: true, // edit before compacting
compactEvery: 1, // compaction hysteresis (turns)
// summarize: (historyText) => ..., // compaction only happens if you provide this
});
Structured results
Give run a resultSchema and the model submits its result through a hidden submit_result tool; the framework validates it and attaches it to result.typed — no more guessing JSON out of plain text.
const { result } = await app.run(messages, {
resultSchema: {
type: 'object',
properties: { pass: { type: 'boolean' }, reason: { type: 'string' } },
required: ['pass', 'reason'],
additionalProperties: false,
},
});
console.log(result.typed); // { pass: true, reason: '…' }The schema can also come from zod (an optional peer): fromZod(z.toJSONSchema(S), S). Validation runs through zod, and the error path goes back to the model verbatim so it can correct itself.
Observability: trace as a first-class citizen
One run == one trace (traceId === runId), built in from turn 0 — not a third-party tracing SDK bolted on. This is the line between Agentia and an agent demo that stops at “it runs”: the four kinds of capability decide what it can do, the trace decides whether you dare to ship it.
- Every step accounted for: each span carries model, input / output / cache tokens, estimated cost, status and error class. The rule is explicit —
totalUsageaccumulates fromllm.turnonly, while a capability span's usage is the aggregate of its descendants, never double-counted. - The exit is one seam:
TraceSink { export(trace) }— and the run delivers on both the success and failure paths, with a throwing sink unable to affect the run. Persistence, querying, sampling and redaction are all done outside the seam by composing sinks; the framework does not choose your policy (working code inexamples/observability/, write-up indocs/observability.md). - Cost is controllable:
priceOverridespatches the price table; when a model is not in the table the framework emits an explicitusage.unpricedevent (it does not silently treat it as zero); see “Hard cost control” below. - Replayable:
traceToMessages(trace)linearises a finished trace back into messages for the model, as the basis for debugging and for “change the prompt and run it again”. - Exportable:
createOtlpExporter({ endpoint })pushes OTLP and aligns with the OTel GenAI semconv on export (it addsgen_ai.*keys: the run root asinvoke_agent,gen_ai.request.modelonllm.turn, andscoreevents translated togen_ai.evaluation.result, while the existingusage.*keys stay);metricsSink()puts metrics through the same seam too (Prometheus text or OTLP metrics), with zero dependencies throughout. Wiring is in “Scenario guide · Put monitoring on a run”. - Visible locally:
agentia devstarts a local inspector panel sharing one renderer with@migor/trace-view— the site Playground uses the same one, so three surfaces run one implementation and cannot drift.
import { createOtlpExporter } from '@migor/agentia';
const { result } = await app.run(messages);
result.trace.spans; // the call tree: llm.turn / capability spans / events
result.trace.totalUsage; // token totals (accumulated from llm.turn only)
// The exit is one seam: persistence / sampling / redaction / metrics are composed outside it; no policy is built in
await createOtlpExporter({ endpoint: 'http://localhost:4318' }).export(result.trace);Metrics and the tuning report (“which capability is slow / expensive / error-prone”) are in “Ecosystem & observability” below.
Hosting & exports
The server.ts below wires the HTTP host, a SQLite task store and the OTLP exporter together.
import { createServer } from 'node:http';
import { createHttpHandler, SqliteTaskStore, AsyncRunner, createOtlpExporter } from '@migor/agentia';
const runner = new AsyncRunner(app, { store: new SqliteTaskStore('./tasks.db') });
createServer(createHttpHandler(app, { runner })).listen(8080);
// POST /run (sync / SSE) POST /tasks (async) GET /tasks/:id (polling) GET /healthz (probe)
const otlp = createOtlpExporter({ endpoint: 'http://localhost:4318' });
await otlp.export(result.trace);Models & memory
The engine depends only on the ModelClient structural surface: Anthropic satisfies it natively, and OpenAI-compatible endpoints (DeepSeek and friends) plug in with one line of createOpenAIClient; the memory option hydrates and writes back the blackboard between runs.
import { createOpenAIClient, InMemoryMemoryStore } from '@migor/agentia';
const { result } = await app.run(messages, {
client: createOpenAIClient({ baseURL: 'https://api.deepseek.com' }),
model: 'deepseek-chat',
memory: { store: new InMemoryMemoryStore(), keys: ['profile'] },
});Stability & streaming
Cancellation, retries and streaming share one per-call options object, carried from the entry point all the way to the model request.
- Cancellation propagation:
signaltravels all the way toModelClient.messages.stream(both built-in Anthropic and OpenAI adapters forward it); an interruption does not bubble an exception, and the run closes withstopReason='aborted'. The HTTP host builds in “client disconnect ⇒ abort”, andAsyncRunner.runTimeoutMsaborts an in-flight run when it hits. - Retries and backoff: on by default (
maxAttempts=3, exponential backoff + jitter), and only failures where the attempt produced no text at all are retried — text already streamed cannot be taken back. Each attempt opens its ownllm.turnspan. - SSE streaming:
POST /runwithAccept: text/event-streamemitstext.delta/run.end/errorframe by frame; without that header it still answers with unary JSON. With backpressure protection: once downstream backlog (res.writableLength) exceedssseMaxBufferedBytes(8 MiB by default) the stream is closed and the corresponding run is aborted — otherwise one client that connects but never reads would exhaust memory.
const { result } = await app.run(messages, {
signal: controller.signal, // cancellation propagates into in-flight model requests
retry: { maxAttempts: 3, baseDelayMs: 500 }, // on by default; retry: false turns it off
onText: (delta) => process.stdout.write(delta), // streaming text deltas
});Host hardening
The framework does not implement token / JWT auth, and it does not subscribe to process signals — that belongs to the host or a reverse proxy.
authenticateintercepts at the entry, before the body is read (including unknown paths, so it never reveals whether a path exists); throwingHttpExceptionanswers with its status / body, while any other throw answers 401 with the detail going only to server logs.drain({ timeoutMs }): refuse new work → wait for in-flight work to finish → force-close any SSE streams still open at the timeout; unfinished tasks stay in the store and the next start callsresumePending()to continue them (they are not discarded).GET /healthz→{ ok, inFlight, uptimeMs, draining }: unauthenticated (a probe cannot carry credentials) and still answering 200 while draining;inFlight= in-flight synchronous runs + queued/running async tasks.
const handler = createHttpHandler(app, {
authenticate: async (req) => {
if (!verify(req.headers.authorization)) throw new HttpException(401, { error: 'invalid credentials' });
},
});
createServer(handler).listen(8080);
process.on('SIGTERM', () => handler.drain()); // signal handling belongs to the hostHard cost control
This is complementary to the long-context policy, not the same thing: the context policy changes the messages before sending (to avoid 400s / premature compaction), whereas this changes the run's outcome after accounting — its purpose is controlling spend.
maxTotalTokens/maxCostUsd: checked after each turn's accounting; going over closes the run withstopReason='budget_exceeded'(counted as a failure) and records abudget.exceededevent on the run root. The scope is the whole run (sub-agents included).- Two deliberate semantics: a turn where the model finished naturally is not re-labelled a failure for going over (the task did in fact complete); and on overrun the framework does not execute that turn's tools (so no further side effects are produced).
toolTimeoutMs: a timeout does not kill the run — that tool_result is recorded asis_errorso the model takes another route;maxToolConcurrencycaps parallel tools within one turn. ⚠️ A timeout means “we stopped waiting”, not cancellation —AgentTool.runreceives no signal. It is the sole arbiter: when it is set, a tool's own timeout (the MCP bridge'stimeoutMs) takes no part in the decision.priceOverrides: patches the price table — override a built-in model's rates, or price a non-Anthropic endpoint (DeepSeek / self-hosted). It is the precondition formaxCostUsdto work at all: with an unlisted model the cost stays 0 and the guardrail never fires, so the framework says so explicitly via theusage.unpricedevent, themodel_unpriced_turns_totalmetric and theonUnpricedModelcallback — never silently.
const { result } = await app.run(messages, {
maxTotalTokens: 200_000, // cumulative token cap (whichever comes first)
maxCostUsd: 0.5, // cumulative cost cap (requires the model to be in the price table)
toolTimeoutMs: 10_000, // per-tool timeout → is_error, does not kill the run
maxToolConcurrency: 4, // cap on parallel tools within one turn
});Ecosystem & observability
All three go through existing extension points: an MCP tool is just an AgentTool, metrics are just a TraceSink, and evals are assertions written against a trace.
- MCP:
mcpTools(client, { server })maps a server'stools/listonto menu entries. The connectors ship in the box (createStdioMcpConnector/createStreamableHttpMcpConnector, using onlynode:child_process+ globalfetch); the framework does not import the MCP SDK, and theMcpClientLikeseam remains (the official SDK / a remote server / your own transport all still fit). Characters that are unfriendly to an LLM are normalised to_, and the original name is recorded as themcp.tool.<menu name>attribute keyed by menu name (one entry per call, so parallel calls do not overwrite each other) for auditing / replay. There is one timeout arbiter: when the engine setstoolTimeoutMs, the bridge'stimeoutMstakes no part in the decision (a connector's owntimeoutMscovers only the assembly-time handshake andtools/list), and both paths share one verdict and one accounting entry (errorKind='timeout'). - evals:
scriptedClient([...])scripts the model round-trips anddefineEval({ app, cases, expect })writes assertions against the trace — whether a prompt change, a model swap or a new tool caused a regression is answered by assertions rather than by eye. Failing cases carry the full trace. - Metrics:
metricsSink()satisfiesTraceSinknatively, and all three dimensions are derived from the trace with nothing to instrument — run level, plus capability level (call count, failure count, latency, tokens, cost for each tool / skill / subagent: tools read thetool.outputevent on the turn, skills / subagents read the capability span) and model level (attributed per model, plus a count of “turns whose cost could not be computed”). Durations come in two flavours: a histogram that aggregates across instances, and exact percentiles within the window;render()produces Prometheus text,createHttpHandler({ metrics })builds inGET /metrics, andexport:'otlp'pushes to an OTel collector (hand-written JSON, zero dependencies). Score-level metrics:attachScorewrites ascoreevent onto the run root which aggregates intoagentia_score{name,source}(gauge, latest value) andagentia_score_total(counter, number of entries) — eval verdicts / LLM-judge / human labels all take this route; a ready-made Grafana dashboard ships with the repo (examples/observability/grafana-dashboard.json). Note thatcreateHttpHandler({ metrics })only renders — for the numbers to actually accumulate you must wire the same sink intocreateApp({ sinks: [metricsSink()] }), otherwise they stay 0 and nothing is reported. - Tuning report:
buildRunReport(trace)turns a trace into a ranking of which capability is slow / expensive / error-prone (mergeRunReportsaggregates across runs), and one CLI lineagentia report traces.jsonlprints the table (you produce that jsonl yourself: attach aTraceSinkthat appends JSON lines, or export fromFileTaskStore); theagentia devpanel shows the same ranking. Faced with a pile of knobs, this is your basis for deciding which to turn. - Prompt versioning:
new SystemPrompt({ version })writes the version into the run root assystem.version, so a trace can answer “which version of the prompt produced this result”. Multi-tenant quotas: combinemiddleware(intercept) +TraceSink(accounting) + BudgetGuard (per-call cap).
import { createApp, createStdioMcpConnector, mcpTools, metricsSink, defineEval, scriptedClient } from '@migor/agentia';
// 1) MCP: a server's tools become menu entries (the connector ships in the box; you can also pass any object implementing listTools/callTool)
const mcp = createStdioMcpConnector(['uvx', 'mcp-server-time']);
const tools = await mcpTools(mcp, { server: 'time' });
const app = createApp({ system, providers, tools, sinks: [metricsSink()] });
// 2) evals: scripted round-trips + assertions against the trace
const report = await defineEval({
name: 'time-flow',
app: () => app,
cases: [{ input: 'What time is it in Shanghai?', client: scriptedClient([toolCallTurn, finalTurn]) }],
expect: (r, { trace }) => { /* assert stopReason / which tools were called / typed */ },
}).run();
if (!report.ok) console.error(report.cases.filter((c) => !c.ok));
// 3) metrics: Prometheus text; expose a route that returns metricsSink().render()Approval, guardrails & code isolation
For guardrails and code isolation the framework only gives the seam, not a subsystem (the policies differ too much for hard-coding to be right); approval, beyond the gate seam, also has built-in suspend / resume (HITL). Every recipe uses existing extension points only; for the full boundaries see the Known boundaries section and /llms-full.txt.
Approval gate
Middleware can await a human decision before letting a call through — the engine will wait for the tool result. Allow = next(); reject = do not call next() (short-circuit, the side effect never happens); reject and redirect the model = throw (recorded as is_error, run continues). For durable approval that survives a process or restart, use @Tool({ approval: 'required' }): the run suspends as suspended (message history persisted), and POST /tasks/:id/approve resumes execution with that decision — see AsyncRunner.approve and ApprovalDecision in the API reference.
import { createApp } from '@migor/agentia';
import type { CapabilityMiddleware } from '@migor/agentia';
const DANGEROUS = new Set(['send_email', 'deploy', 'delete_records']);
const requireApproval: CapabilityMiddleware = async (call, next) => {
if (!DANGEROUS.has(call.capability.name)) return next();
const ok = await askHuman(call.capability.name, call.input); // wait in place (the run stops here)
if (!ok) throw new Error(`not approved: ${call.capability.name}`); // → is_error, the model reroutes
return next(); // allow
};
createApp({ system, providers, middleware: [requireApproval] });Content guardrails
Input (wrap app.run; on the HTTP side use createHttpHandler({ authenticate }), which intercepts before the body is read), before a tool call (middleware), and output (wrap the return value, or attach sinks).
const callable = {
name: app.name,
async run(msgs, opts) {
const out = await app.run(redactInput(msgs), opts); // input guardrail
if (flagged(out.result.finalText)) throw new Error('output blocked by guardrail'); // output guardrail
return out;
},
};Code execution isolation
The framework never executes model-generated code — @Skill runs a method body you wrote and @Tool runs a function you wrote, so model output only ever becomes text / a tool_result. Isolation is done inside the tool implementation (Docker / a subprocess / a micro-VM, your call) and the framework takes no part in it — the AgentTool.run(input) → output contract keeps isolation entirely within the implementation. The one place it touches the edge: if a tool spawns a subprocess, it must kill it on cancellation (the framework's timeout means “stopped waiting”, not cancellation).
Known boundaries (stated plainly)
What the framework does not guarantee — the basis for deciding whether you can use it. It is a long list (dozens of entries, each one guarded in the repo), and it lives in a single source: docs/usage-guide.md §7, which is written in Chinese. That same text is served verbatim at /llms-full.txt, and the Chinese docs page renders the whole table at build time.
Read it at the source rather than trusting a copy here: docs/usage-guide.md §7 on GitHub · /llms-full.txt. The English summary of the same commitments is under Versioning & stability below.
CLI reference
Install globally with npm i -g @migor/cli and use the agentia command.
| Command | What it does |
|---|---|
| agentia create <name> | Scaffolds a new project: the four category folders (with .gitkeep), a src/registry.ts registry, src/app.ts (the assembly factory) + src/main.ts (a thin entry point) + src/dev.config.ts, and tsconfig. |
| agentia g <type> <name> | Generates a capability into the matching category folder (tool → src/tools/, skill → src/skills/, prompt → src/prompts/, subagent → src/subagents/) and registers it in src/registry.ts automatically; long-text assets (system.md / asset.md) are generated along with it. |
| agentia dev | Starts a local inspector panel: type a prompt to drive one real run (with ↑/↓ history), multi-select capabilities (narrowing the menu), choose the working directory, toggle multi-turn, then inspect the call tree in trace-view. The CLI manages file watching itself (the allow-list includes .md, so editing a text asset needs no restart). Binds 127.0.0.1 only, with Origin checks and a one-time token per launch. |
| agentia doctor | Assembly health check: unregistered / dangling capabilities, naming rules, duplicate entries. |
| agentia report <trace.jsonl> | Builds a tuning report from an on-disk trace file: capability latency / cost / error-rate rankings (merged per capability across lines). One JSON object per line — either a bare Trace or a TaskRecord containing result.trace. You produce the file: attach a TraceSink that appends JSON lines, or export from FileTaskStore. |
| agentia harvest <trace.jsonl> | Turns a production trace into an eval case scaffold (--out <file.ts> / --failed / --limit N): the scriptedClient script is rebuilt from the main loop's llm.turn entries, and expect is pre-filled with tool-sequence assertions. A scaffold is not a finished test — assistant text is a placeholder (the trace records no text), so review by hand before it enters CI. |
| agentia diff <a.jsonl> <b.jsonl> | A/B comparison of the call trees in two on-disk trace files (same input shape as agentia report): prints a run-level summary plus per-span differences, and exits 1 when they differ — so CI can block “the trajectory drifted after a prompt / model change”. The in-code equivalent is diffTraces. |
| agentia add <pkg> | Installs a third-party capability package (a defineModule package) and registers it in the registry. |
Environment variables
| Variable | Notes |
|---|---|
| ANTHROPIC_API_KEY | Required. Anthropic API key (ANTHROPIC_AUTH_TOKEN / ANTHROPIC_BASE_URL are also supported for compatible endpoints). |
.env (loadEnvFile()) | The framework does not read .env on its own: the scaffold's loadEnvFile() in src/app.ts loads it into process.env (it lives in the assembly module so that both entry points — agentia dev and npm start — pick it up). Real environment variables win — a key that is already defined (even as an empty string) is not overwritten by the file unless you pass { override: true } explicitly. |
| AGENTIA_MODEL | Optional. Global default-model override; falls back to claude-opus-5 when unset. An explicitly passed model wins over everything. |
| OPENAI_API_KEY | Optional. Key for OpenAI-compatible providers, used in multi-provider setups. |
Versioning & stability
What 0.x means: version numbers follow SemVer, but during
0.x a minor release may contain breaking changes. There is exactly
one rule — a breaking change must leave a trace: the corresponding entry in
CHANGELOG.md must carry a dedicated section for it (with a heading containing the
migration wording) that spells out what you have to change.
# List every section that asks you to act (this page deliberately hard-codes no count — written numbers rot)
grep -nE '^#{3,4} .*(迁移|破坏性)' CHANGELOG.md
⚠️ Stated plainly: this discipline was tightened over time — an
early 0.7.1 (a type-surface change) put the migration step in a block quote at the
top of the version entry with no dedicated section. So that command gives a lower
bound, not the full set; to read everything, read the whole entry for a release.
| Surface | Promised | Not promised |
|---|---|---|
| Public API | The named exports of @migor/agentia (source of truth src/index.ts, with a reverse full-coverage guard on the site's API page) plus the command-line surface of @migor/cli | Deep-path imports (dist/**), internal module paths, packages/*, examples/* — this boundary is sealed by the exports field in package.json (guarded by tests/architecture/package-exports.test.ts) |
| Behaviour | The migration steps written in that version's entry | Any behaviour change not written into the migration notes |
| Runtime | The range declared in engines.node; CI runs the whole chain on every version in that range | Deno / Bun / edge runtime (unverified); EOL Node versions (dropped in a minor) |
| Dependencies | Zero third-party runtime dependencies (guarded by tests/architecture/no-runtime-deps.test.ts, not a verbal promise) | devDependencies and the dependency surface of examples/* |
The 1.0 bar
All three are checkable, not slogans:
- Every item in
docs/spec.md§11 “open items” is either delivered or explicitly struck out (with a written reason why it does not block); docs/guards.md§2 “to be guarded” is empty — every declared boundary is held by a machine rather than by someone remembering;- The public export surface goes three consecutive minors without a breaking change.
Performance order of magnitude
The framework ships five benchmarks (scripts/bench-*.ts). They answer
shape questions — which cost is paid under what conditions, and what a different
implementation would save — not an SLA.
| Your question | Local reading (order of magnitude) | Shape conclusion (machine-independent) | Re-run |
|---|---|---|---|
| How expensive is assembling an app | Pure assembly 0.4 ms; discovery + assembly 8.1 ms; whole-process start + assembly 4040 ms (of which cold assembly ≈2963 ms, and process start alone is 1077 ms) | In a long-lived runner, changing the capability selection pays only S2 (≈8 ms) and not S6b (≈4 s) ⇒ a ~500× difference; of those 4 s, ~1.1 s is tsx starting up, unrelated to the framework | npm run bench:assembly |
| When does an MCP connector cost you | Cold 279.5 ms; warm (connector reused) 0.3 ms | Only rebuilding the connector pays the cold cost ⇒ a ~900× difference | npm run bench:assembly |
| Does recording trace content cost much | Large outputs (≈37 KB each) over 5 calls: default 17.2 KB / untruncated 233.9 KB (13.6×) / truncated at 200 6.2 KB | The default already truncates; traceContent:'full' is a switch that charges by output bytes (measured at only 1.0× for small outputs) | PAYLOAD_ROWS=1000 CALLS='[5,20]' npm run bench:trace |
| Cost of a metrics snapshot / OTLP flush | snapshot() 6.56 ms; buildOtlpPayload() 0.073 ms | Flush cost is almost entirely in snapshot() ⇒ skipping it saves ≈6.5 ms per call, proportional to flush frequency | npm run bench:otlp |
| Due-scanning of suspended tasks | sqlite N=10000: full-table list() 61.9 ms / due index listDue 0.1 ms | Whether the index exists is structural ⇒ a ~600× difference, not a tuning matter | npx tsx scripts/bench-resume-scan.ts 10000 |
Shape of the redis store's list() | N=10000: 10001 round-trips, 42.1 MB parsed, pure CPU 104.1 ms; with an auxiliary ZSET ⇒ 1001 round-trips | Round-trip count is a structural quantity (the thing a different implementation would save); latency is extrapolated only (no redis instance on this machine) | npx tsx scripts/bench-redis-due.ts 10000 |
⚠️ Milliseconds are not a promise: they move with the machine / Node version / load, and copying them as an SLA would be wrong. What stays constant is the shape (how many times larger, how many round-trips, what a given cost is proportional to). The five benchmarks deliberately stay out of CI (timing benchmarks only add noise on CI machines) ⇒ nothing guards these numbers, so if you want them, run them yourself using the “Re-run” column.