astorlm v0.2.2
npm
SDK-first · no CLI, no TUI

The agent loop,
as a library you embed.

astorlm gives you the moving parts of a coding agent — the multi-turn loop, Zod-typed tools, MCP, skills, subagents, sandboxed execution, tracing and evals — and stays out of your way. You import it and build your own harness on top.

$ pnpm add astorlm
🧩

Small conceptual surface

Provider, tools, agent, loop, executor. Five pieces you can hold in your head.

🌍

Runtime-agnostic core

Node, Deno, Bun, edge and browser. Anything needing Node sits behind its own subpath.

🔌

MCP, both ways

Mount MCP servers as tools, including MCP Apps interactive UI (SEP-1865).

🛡️

Isolation you choose

Local, Docker, or a capability-empty WASM sandbox for code the model writes.

📊

Observable by design

OpenTelemetry spans, cost metrics, deterministic replay and an eval harness.

🪫

Built for weak models

edge-boost hardens the loop for free tiers, small local models and quantized checkpoints.

Introduction

astorlm is an embeddable agentic library for TypeScript. SDK-first: there is no CLI and no TUI — you import it and build your own harness on top.

The conceptual model is small:

  • Provider — a model client (Anthropic, OpenAI-compatible) that streams events.
  • Tools — Zod-typed functions the model can invoke.
  • Agent — the returned instance holding in-memory state (messages, registry, event bus).
  • Loop — the multi-turn cycle that wires model + tools + hooks together.
  • Executor — the backend that runs commands (local, docker, custom).

Designed to be runtime-agnostic in the core (Node, Deno, Bun, edge, browser), with a separate subpath for anything that requires Node (reachable under astorlm/core).

Installation

pnpm add astorlm
# or
npm install astorlm

Requirements: Node ≥ 20, TypeScript with ESM resolution ("type": "module" in your package.json).

Quick start

A coding agent with built-in tools, pointed at a local OpenAI-compatible endpoint:

import { OpenAIProvider, createLocalAgent } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'

const agent = await createLocalAgent({
  cwd: process.cwd(),
  provider: new OpenAIProvider({
    model: 'myModel',
    baseURL: 'http://127.0.0.1:11434/v1',
    apiKey: 'myApiKey',
  }),
  tools: createCodingTools(),
})

agent.on('text', (text) => process.stdout.write(text))

await agent.run('List the .ts files in this project.')
console.log(agent.getUsage())
Note. createAgent and createLocalAgent are async — always await them. Session manager state is loaded before they return.

Example conventions

Every code sample in this wiki uses the same placeholders, so there is only one thing to swap:

PlaceholderStands for
'myModel'The chat model id your endpoint serves.
'myEmbeddingModel'The embedding model id (embeddings samples only).
'myApiKey'Your API key. Purely local endpoints usually ignore it, but it must be a non-empty string.
'http://127.0.0.1:11434/v1'Any OpenAI-compatible base URL. Swap it for your own.

Substitute your own values, or read them from the environment — OpenAIProvider can pick the key up from an env var for you (see Providers).

Entry points

ImportRuntimeWhat it exposes
astorlmagnostic / NodeCore plus Node re-exports: createAgent, createLocalAgent, OpenAIProvider, AnthropicProvider, tool, ToolRegistry, FileSessionManager, AstorAgent, createSubagentTool, runGoalLoop, generateObject, etc.
astorlm/coreNode ≥ 20createLocalAgent, FileSessionManager, AstorAgent, LocalExecutor, DockerExecutor, mountMcpServer, createFileSystemSkillSource, createLayeredSkillSource
astorlm/toolsNode ≥ 20createCodingTools, createReadOnlyTools, individual built-in tools (formerly astorlm/tools/node)
astorlm/promptagnosticbuildSystemPrompt, DEFAULT_SYSTEM_PROMPT, compilePrompts, formatPromptReport — see Prompt compiler.
astorlm/embeddingsagnostic (fetch)createOpenAIEmbedder, createSemanticIndex, withEmbeddingCache, cosineSimilarity / dotProduct / euclideanDistance — also re-exported from the astorlm barrel.
astorlm/experimental/contractNode ≥ 20createContractHooks, ContractViolationError and the contract types.
astorlm/experimental/error-registryNode ≥ 20createErrorRegistry, errorRegistryHooks
astorlm/experimental/tracingagnosticattachTracer, createTracer, createInMemoryExporter and span types (observability).
astorlm/experimental/tracing/otelagnostic (fetch)createOtlpSpanExporter — OTLP/HTTP export to OpenTelemetry collectors.
astorlm/experimental/metricsagnosticattachMetrics, computeCost, resolvePricing — operational metrics and USD cost.
astorlm/experimental/replayagnosticcreateRecordingProvider, createReplayProvider — deterministic record & replay.
astorlm/experimental/evalsagnosticrunEval plus scorers (exactMatch, contains, regexMatch, toolTrajectory, llmJudge).
astorlm/experimental/wasm-runneragnostic (WASM)QuickJsCodeRunner, createCodeRunnerTool — capability-empty code sandbox. See WASM code sandbox.
astorlm/experimental/edge-boostagnosticedgeBoost, createGuardedProvider, mergeSessionHooks — hardening profile for weak, local or free models.

createAgent

The core factory. Returns an in-memory Agent with no Node dependencies.

Signature

createAgent(opts: CreateAgentOptions): Promise<Agent>

Options

FieldTypeDescription
providerProviderModel client. Required.
cwdstringLogical working directory for the session.
toolsTool[]Tools registered up front.
systemPromptstringReplaces the default system prompt.
appendSystemPromptstringExtra text appended at the end of the prompt.
maxTurnsnumberTurn cap per run(). Default 25.
hooksSessionHooksSee Hooks.
contextOptimizerContextOptimizerOptions | falseAutomatic compaction. See optimizer.
retryRetryPolicyRetries for transient errors.
executorExecutorBackend for bash. Default: noop.
sessionManagerSessionManagerPersistence + branching.
sessionIdstringResumes an existing session from the manager.
skillSourcesSkillSource[]Loadable skill sources.
skillMode'filesystem' | 'on-demand' | 'all'How skills are exposed to the model.
pattern'REACT' | 'PLAN_EXECUTE'Loop style. Default 'REACT'. See Loop patterns & Heartbeat.
heartbeatHeartbeatOptionsConfig for the proactive background loop. See Heartbeat.

Agent members

MemberDescription
id: stringUnique session id.
pattern: AgentLoopPatternLoop style the agent was created with ('REACT' or 'PLAN_EXECUTE').
run(input, { abortSignal? })Pushes a user message (string or structured) and runs the loop to completion. Returns the final response.
fork({ newSessionId?, branchFromMessageId? })Creates a branched child session and returns a new initialized Agent.
on(event, listener)Registers a listener for a specific event ('text', 'thinking', 'tool-start', 'tool-end', 'error', 'event'). Returns an unsubscribe function.
abort(reason?)Cancels the in-flight turn.
getMessages()Snapshot of the history.
getUsage()Accumulated token usage.
registerTool(t)Adds a tool at runtime.
startHeartbeat(opts?)Starts the proactive background loop. Accepts HeartbeatOptions to override the initial config.
stopHeartbeat()Stops the active heartbeat (clears the interval and the global timeout).
getPlan()Snapshot of the current PlanItem[] (in PLAN_EXECUTE mode).

createLocalAgent node

A wrapper over createAgent that automatically injects a LocalExecutor (so the bash tools work) and a fileReader backed by node:fs (so the system prompt can pick up AGENTS.md / CLAUDE.md when they exist).

Same options as createAgent. Use it whenever you are on Node.

AstorAgent node

An ergonomic facade over createLocalAgent for common flows (one prompt, a fresh session per task, branching).

import { AstorAgent, FileSessionManager, OpenAIProvider } from 'astorlm'

const agent = new AstorAgent({
  provider: new OpenAIProvider({
    model: 'myModel',
    baseURL: 'http://127.0.0.1:11434/v1',
    apiKey: 'myApiKey',
  }),
  sessionManager: new FileSessionManager({ dir: './.astor-sessions' }),
  defaultOutputMode: 'verbose',
})

const { sessionId, text } = await agent.run('Add 2 + 2')

const childAgent = await agent.fork({
  parentId: sessionId,
})
await childAgent.run('Now subtract them')

Output modes

  • 'silent' — subscribes to nothing.
  • 'console' — streams text to stdout.
  • 'verbose' — adds tool execution and session end logs.
  • (event) => void — custom function: you receive the raw AgentEvent.

Events

Register listeners through the session's session.on(event, listener) method:

EventListener argumentDescription
'text'text: stringEach chunk of the model's response text, in real time.
'thinking'thinking: stringEach chunk of reasoning content, in real time.
'tool-start'{ name, input }Fires before a tool runs.
'tool-end'{ name, output, isError }Fires after a tool finishes.
'error'error: anyFires if the session ends with an error.
'event'event: AgentEventThe raw telemetry event off the bus (e.g. turn_start, session_end).

Among the raw events reachable through 'event', the loop variants emit:

Raw eventPayloadDescription
heartbeat_tick{ checkPrompt }The heartbeat woke up and is about to run (it already passed localCondition, if one was set).
plan_updated{ plan: PlanItem[] }The plan changed after an add_plan_item / update_plan_item call (PLAN_EXECUTE mode).
user_steering{ toolUseId, feedback }A hook steered the agent at a tool boundary. See Steering.
provider_retry{ attempt }A transient provider error is being retried. See Retry.
contract_violation{ kind, detail }The agent contract was breached. See Agent contract.

Providers

An adapter over a model. astorlm ships two out of the box.

OpenAIProvider

Chat Completions with tool_calls. Works against any OpenAI-compatible endpoint: Groq, OpenRouter, Together, Ollama, vLLM, llama.cpp.

new OpenAIProvider({
  model: 'myModel',
  baseURL: 'http://127.0.0.1:11434/v1', // optional, defaults to api.openai.com
  apiKey: 'myApiKey',                    // or read from an environment variable
  envVar: 'OPENAI_API_KEY',                // name of the env var (e.g. GROQ_API_KEY)
  maxTokens: 4096,
})
Free tiers, local models, and not shipping API keys around. Because baseURL accepts any OpenAI-compatible endpoint, the cleanest way to run astorlm against free tiers, local models, or a mix of several backends is to point it at a proxy instead of at a vendor directly. OperatorLM is a companion project built for exactly that: it exposes one local OpenAI-compatible endpoint, holds the real provider keys on its own side, and routes each request across several targets with failover — so your agent code carries a placeholder key like 'myApiKey' and never sees a real credential. Recommended whenever you want free/local models, key isolation, or per-model routing without touching the agent.

Note that model failover between providers is deliberately out of scope for astorlm — that belongs to the routing layer. What astorlm owns is what happens once a weak model is already answering badly: see edge-boost.

AnthropicProvider

Native streaming through @anthropic-ai/sdk. Supports thinking blocks and reports cache read/creation.

new AnthropicProvider({
  model: 'claude-3-5-sonnet-latest',
  apiKey: process.env.ANTHROPIC_API_KEY,
  maxTokens: 4096,
  baseURL: 'https://api.anthropic.com', // optional
})

Custom provider

Implement the Provider interface (a single stream(opts) method emitting ProviderEvents) and pass it to the session. Useful for mocking in tests, or for adapting SDKs that are not OpenAI-compatible.


Embeddings & semantic search

First-class embedding primitives (representing text as dense vectors so you can compare by meaning rather than by literal match): they live next to the providers in the astorlm barrel and also under the astorlm/embeddings subpath. They are runtime-agnostic (only fetch, zero dependencies) — they run wherever fetch runs: Node, Deno, browser, edge.

createOpenAIEmbedder

An OpenAI-compatible client. embed (one text) and embedMany (batched, a single round-trip), both reporting token usage. Works against OpenAI, Ollama, vLLM, LM Studio and friends by changing baseURL.

import { createOpenAIEmbedder } from 'astorlm'

const embedder = createOpenAIEmbedder({
  baseURL: 'http://127.0.0.1:11434/v1',
  model: 'myEmbeddingModel',
  apiKey: 'myApiKey',
  // dimensions: 256,   // optional (text-embedding-3-* models truncate)
})

const { embedding, usage } = await embedder.embed('hello world')
const { embeddings } = await embedder.embedMany(['a', 'b', 'c']) // 1 request

createSemanticIndex

An in-memory vector store (exact brute-force search). This is the reusable primitive behind semantic search, RAG, and the error registry's matching. You index text and query it by meaning with topK (how many results) and threshold (minimum score).

import { createOpenAIEmbedder, createSemanticIndex } from 'astorlm'

const index = createSemanticIndex({ embedder })

await index.addMany([
  { id: 'oom', text: 'The Node process dies with out of memory during the build.' },
  { id: 'tls', text: 'TLS handshake fails: expired certificate.' },
])

const hits = await index.query('it ran out of RAM while compiling', { topK: 1 })
// → [{ id: 'oom', score: 0.8…, text: '…' }]

Other methods: add, addVector (pre-computed vector), queryByVector, remove, get, list, clear, size. Upsert by id. Generic over the metadata type (createSemanticIndex<M>).

withEmbeddingCache

Wraps any Embedder and memoizes identical inputs (no re-embedding, no re-billing). In embedMany it asks the backend for only the cache misses, in a single batch.

import { withEmbeddingCache } from 'astorlm'

const cached = withEmbeddingCache(embedder)

Similarity helpers

cosineSimilarity returns the standard [-1..1] cosine (1 = identical, 0 = orthogonal, -1 = opposite). Also dotProduct and euclideanDistance. All of them return 0 for mismatched-length or empty vectors instead of throwing.

Back-compat. The cosineSimilarity re-exported by astorlm/experimental/error-registry stays clamped to [0..1] (so it is comparable to the Jaccard fallback); the first-class one in astorlm returns [-1..1]. createOpenAIEmbeddingClient still exists as a deprecated alias.

Tools

A tool is a Zod-typed function the model can invoke. You define them with tool.

tool

import { tool } from 'astorlm'
import { z } from 'zod'

const getWeather = tool({
  name: 'get_weather',
  description: 'Returns the current weather for a city.',
  schema: z.object({
    city: z.string().describe('City name'),
    units: z.enum(['c', 'f']).default('c'),
  }),
  async execute({ city, units }, ctx) {
    // ctx: { cwd, abortSignal, logger, executor }
    return `In ${city}: 22°${units}`
  },
})

Built-in tools node

NameArgumentsNotes
readpath, offset?, limit?1-indexed line numbering. Default 2000 lines.
writepath, contentOverwrites.
editpath, oldString, newString, replaceAll?oldString must be unique unless replaceAll.
bashcommand, timeoutMs?Synchronous. Default 120s, max 10min.
bash_spawncommandBackground. Returns a pid.
bash_get_outputpidDrains the buffer.
bash_killpid, signal?Sends a signal to the process.
lspath?Directory listing.
grepregex patternContent search.
globpatternPath search.

Helpers

import { createCodingTools, createReadOnlyTools } from 'astorlm/tools'

createCodingTools()    // all of them, including background bash
createReadOnlyTools()  // read + ls + grep + glob

Subagents (agent-as-tool)

createSubagentTool exposes a whole independent agent to a parent agent as a single tool. When the parent calls it, the tool spins up a fresh child session with its own (usually narrower) system prompt and tool set, runs one prompt to completion, and returns the child's final text as the tool result. The parent never sees the child's intermediate turns — only the distilled answer, which is what keeps the parent's context window from filling up with the child's exploration.

This is pure composition over the public API: it does not touch the loop or the session internals. The child inherits the parent's cwd, executor and logger from the ToolContext, and the parent's abort signal is propagated, so cancelling the parent cancels the child mid-flight.

import { OpenAIProvider, createLocalAgent, createSubagentTool } from 'astorlm'
import { createCodingTools, createReadOnlyTools } from 'astorlm/tools'

const provider = new OpenAIProvider({
  model: 'myModel',
  baseURL: 'http://127.0.0.1:11434/v1',
  apiKey: 'myApiKey',
})

const researcher = createSubagentTool({
  name: 'research_agent',
  description: 'Delegate a read-only investigation of the codebase. Returns a summary.',
  provider,                                // may differ from the parent's — e.g. a cheaper model
  appendSystemPrompt: 'Be concise. Report findings only.',
  tools: createReadOnlyTools(),           // narrower surface than the parent's
  maxTurns: 8,
})

const parent = await createLocalAgent({
  provider,
  tools: [...createCodingTools(), researcher],
})

await parent.run('Find where retries are implemented, then add a test for them.')

Options (SubagentToolOptions)

FieldTypeDescription
namestringTool name the parent model invokes (e.g. 'research_agent'). Required.
descriptionstringWhen the parent should delegate to this subagent. Required.
providerProviderProvider the child session uses. May differ from the parent's. Required.
systemPromptstringSystem prompt for the child. Defaults to the SDK base prompt.
appendSystemPromptstringExtra text appended to the child's system prompt.
toolsTool[]Tools the child may use. Defaults to none (pure reasoning).
maxTurnsnumberTurn cap for the child loop.
pattern'REACT' | 'PLAN_EXECUTE'Loop pattern for the child. Defaults to 'REACT'.
inputKeystringInput property the parent fills with the delegated task. Defaults to 'task'.
inputDescriptionstringDescription of that input property, shown to the parent model.
Why scope the child's tools down. A subagent is the natural place to enforce least privilege: give the researcher read-only tools and the parent keeps the write capability to itself. Combine it with agent contract hooks for a hard budget per delegation.

Structured output (generateObject)

Returns an object that is typed and validated against a Zod schema, instead of plain text. The general primitive is the terminal tool: a synthetic tool is registered whose schema is the desired result, and the model is instructed to call it to deliver the answer; its input is the object. This reuses the agent's tool validation and repair loop, works on any provider, and coexists with other tools.

Case 1 — single-pass extractor (no tools)

With no tools, generateObject picks 'native' mode on its own: it uses the provider's response_format: json_schema (constrained decoding — the model cannot emit tokens outside the schema). A hard guarantee.

import { OpenAIProvider, generateObject } from 'astorlm'
import { z } from 'zod'

const Ticket = z.object({
  intent: z.enum(['bug', 'feature_request', 'billing', 'other']),
  priority: z.enum(['low', 'medium', 'high']),
  summary: z.string(),
})

const { object } = await generateObject({
  provider: new OpenAIProvider({
    model: 'myModel',
    baseURL: 'http://127.0.0.1:11434/v1',
    apiKey: 'myApiKey',
  }),
  schema: Ticket,
  prompt: 'Classify: "I cannot pay, the card throws a 500, this is urgent"',
})

object.priority // 'high' — typed at compile time

Case 2 — agent verdict (with tools)

With tools, response_format does not apply (it does not coexist with tool_choice), so generateObject uses the terminal tool: the agent investigates with its tools and finally calls the output tool, which breaks the loop and returns the validated object.

import { createReadOnlyTools } from 'astorlm/tools'

const Verdict = z.object({
  decision: z.enum(['approve', 'request_changes']),
  issues: z.array(z.object({ severity: z.enum(['low', 'med', 'high']), file: z.string(), message: z.string() })),
  summary: z.string(),
})

const { object } = await generateObject({
  provider, schema: Verdict,
  tools: createReadOnlyTools(),   // the agent explores the repo, then emits the verdict
  prompt: 'Review the .ts files in this project and deliver a verdict.',
})

if (object.decision === 'request_changes') process.exit(1) // CI

Options (GenerateObjectOptions)

OptionTypeNotes
providerProviderRequired.
schemaZodTypeAnyRequired. Types the result.
promptstringRequired.
mode'auto' | 'tool' | 'native'Default 'auto': 'native' when there are no tools and the provider is OpenAI-compatible, otherwise 'tool'.
toolsTool[]Forces 'tool' mode. The agent uses them before answering.
maxTurnsnumber'tool' mode. Default 8 (leaves room for repair).
maxRepairAttemptsnumber'native' mode. Retries if the output fails validation. Default 2.
toolNamestringName of the terminal tool. Default provide_final_answer.
onEvent(e) => voidStream of agent events.

Returns { object, message, sessionId?, mode }. If the model never delivers the answer (it does not call the terminal tool, or the output still fails validation after the retries) it throws GenerateObjectError. Repair loop: in 'tool' mode an invalid input goes back to the model as an error tool_result (via Zod validation); in 'native' mode the validation error is re-injected and the call is retried.


Hooks

Five optional interception points, all of which may be async.

HookWhenWhat for
beforeTurnStart of every turnObserve { turn, messages }.
beforeProviderCallBefore provider.streamMutate { messages, systemPrompt }.
beforeToolExecutionBefore every toolReturn { authorize, mockResult? }. Permission system + mocking.
afterToolExecutionAfter every toolReturn the final output string the model will see.
afterTurnEnd of every turnObserve { turn, lastMessage }.

Example: ask for authorization before dangerous tools

const session = await createLocalAgent({
  provider, tools: createCodingTools(),
  hooks: {
    async beforeToolExecution({ toolName, input }) {
      if (toolName === 'bash') {
        const ok = await askHuman(`Run: ${input.command}?`)
        return { authorize: ok, mockResult: ok ? undefined : 'denied by the user' }
      }
      return { authorize: true }
    },
  },
})
Composing hooks. Several features ship their own SessionHooks (contract, error registry, skill tool policy, edge-boost). Merge them with mergeSessionHooks from astorlm/experimental/edge-boost instead of hand-writing the chaining.

Steering (interactive redirection)

Steering lets you pause the agent loop right before a tool executes and inject direct instructions (feedback) to redirect it. It is particularly useful for correcting the model when it makes a planning mistake or proposes a destructive action.

Unlike a flat denial (which only returns a denial error to the model), steering adds an interactive interruption where the feedback is injected as a genuine text block in the conversation history. The model therefore understands the reasoning and corrects its behaviour naturally on the next turn.

Configuring it in beforeToolExecution

To activate steering, the beforeToolExecution hook must return steer: true alongside the feedback text:

const session = await createLocalAgent({
  provider,
  tools: createCodingTools(),
  hooks: {
    async beforeToolExecution({ toolName, input }) {
      if (toolName === 'write' && input.path === 'DELETE_ME.txt') {
        // Pause and interactively ask the user for feedback
        const feedback = await askHumanForFeedback()
        return {
          authorize: false,
          steer: true,
          feedback: feedback || 'Action cancelled.'
        }
      }
      return { authorize: true }
    }
  }
})

Effects on the execution loop

  • Concurrent cancellation: if the model tried to run several tools in parallel and one of them triggers steering, the rest are automatically cancelled in cascade (returning a descriptive cancellation error) so the state is not polluted.
  • History injection: the tool result message (role: 'user') is built containing the cancelled tool_results plus a final text block formatted as [User Steering Feedback]: <feedback>.
  • Event bus: a { type: 'user_steering', toolUseId, feedback } event is emitted for logging and telemetry.

createSteeringController

Writing that hook by hand is fine for a fixed rule, but for the common case — a human watching a run who wants to interject — createSteeringController packages it: you queue feedback from anywhere (a keypress handler, an HTTP endpoint, a UI button) and it is consumed at the next tool boundary.

import { createSteeringController, createLocalAgent } from 'astorlm'

const steering = createSteeringController()

const agent = await createLocalAgent({
  provider,
  tools: createCodingTools(),
  hooks: steering.hooks,      // pass its hooks straight in
})

// from anywhere else — e.g. a keypress handler while the run is in flight
steering.steer('Stop editing src/, work under test/ instead.')

steering.pending   // the queued feedback, or null
steering.clear()   // drop it if it was not consumed yet
MemberDescription
steer(feedback)Queue feedback to steer the agent at the next tool boundary.
clear()Discard queued feedback that has not been consumed yet.
pendingThe currently queued feedback, or null.
hooksThe SessionHooks to pass into createAgent({ hooks }).
Wrapping existing hooks. createSteeringController(base?) accepts a base SessionHooks: when nothing is queued, beforeToolExecution delegates to the base hook (or authorizes by default), and every other base hook is preserved untouched. Feedback only applies at a turn that actually calls a tool — if the agent answers without tools, it is delivered at the next turn that does.

Loop patterns & Heartbeat

The agent runs a multi-turn loop. The pattern option chooses how it structures that work, and heartbeat adds a proactive loop that wakes up on its own in the background.

Patterns (pattern)

ValueBehaviour
'REACT' (default)Classic Reason+Act: the model reasons, calls tools, observes results and repeats until end_turn or maxTurns. No plan state.
'PLAN_EXECUTE'Structured, TodoWrite-style planning. The agent automatically registers two tools (add_plan_item, update_plan_item) and injects an [Active Plan State] block into the system prompt on every turn so the model keeps the task list in view.
Persistent plan & short ids. PlanItems use sequential ids ("1", "2", ...), which are friendlier for small models that have to repeat the id in update_plan_item than a UUID would be. The plan is stored in SessionState.metadata.plan, so it survives resume and fork (the counter continues from the highest restored id).
const agent = await createLocalAgent({
  provider: new OpenAIProvider({
    model: 'myModel',
    baseURL: 'http://127.0.0.1:11434/v1',
    apiKey: 'myApiKey',
  }),
  tools: createCodingTools(),
  pattern: 'PLAN_EXECUTE',
})

agent.on('event', (e) => {
  if (e.type === 'plan_updated') console.log(agent.getPlan())
})

await agent.run('Plan and create a NOTES.md file with a summary.')

Heartbeat (proactive loop)

The heartbeat is a background loop that wakes the agent every intervalMs without you calling run(). Useful for periodic checks (did a file change? is there new work?). Control it with startHeartbeat() / stopHeartbeat(), or let it auto-start by passing heartbeat in the options (unless autoStart: false).

FieldTypeDescription
intervalMsnumberHow often a tick fires. Required.
checkPromptstringThe prompt sent to the model on each tick. Required.
localCondition(cwd) => boolean | Promise<boolean>Latent heartbeat. When defined, each tick runs it locally and only calls the LLM if it returns true. Zero token cost until the trigger fires.
runTimeoutMsnumberPer-tick timeout: aborts the in-flight run() if it takes longer. Decoupled from intervalMs. Default 60000 (60s). Raise it if your model is slow.
timeoutMsnumberTotal time the heartbeat may run before stopping itself. Default 300000 (5 min). 0 or Infinity disables it.
maxTicksnumberTick cap before stopping automatically.
autoStartbooleanWhen false, it does not start on agent creation even though heartbeat was passed. Default true.
Single-flight. An isRunning guard drops ticks that land while a previous run() (or an earlier tick) is still in progress, so turns never overlap. That is why runTimeoutMs is only a safety net for a hung run — keep it above your model's typical latency so you do not cut valid turns short.
// Latent heartbeat: zero cost until the file shows up.
agent.startHeartbeat({
  intervalMs: 8000,
  checkPrompt: 'If PROACTIVE.txt exists, append a line reading "Verified".',
  localCondition: (cwd) => fs.existsSync(path.join(cwd, 'PROACTIVE.txt')),
  runTimeoutMs: 45000,
  timeoutMs: 70000,
  maxTicks: 3,
})

agent.on('event', (e) => {
  if (e.type === 'heartbeat_tick') console.log('tick:', e.checkPrompt)
})

// ...later
agent.stopHeartbeat()
Clean abort. When a tick exceeds runTimeoutMs (or is cancelled), the turn closes as session_end: aborted, not as error. astorlm recognizes both AbortError (fetch/Anthropic) and APIUserAbortError (OpenAI SDK), so a cooperative cancellation is never reported as a hard failure.

Goal loop (Ralph loop)

runGoalLoop is a goal-pursuit orchestrator: it chases an objective across several iterations until it is met. Unlike the heartbeat — which repeats a prompt against the same session as it accumulates context — the goal loop builds a fresh agent on every iteration (clean context window, empty messages[]). State survives between rounds through the filesystem (progress files, the git tree), not through conversation history. It is the pattern behind Codex /goal: plan → act → test → iterate.

FieldTypeDescription
goalstringThe goal to pursue. By default it is the prompt of every iteration. Required.
createIterationAgent(ctx) => Agent | Promise<Agent>Factory returning a fresh agent per iteration. Calling it again each round is what gives the clean context. Use the same cwd so iterations share on-disk state. Required.
isDone(ctx) => boolean | Promise<boolean>Stop condition, evaluated after each iteration. Keep it cheap and deterministic: run the tests through the executor and check the exit code, or look for a marker file. Required.
maxIterationsnumberMandatory fuse: iteration cap so an isDone that never fires cannot run forever. Default 10.
iterationPrompt(ctx) => stringBuilds the prompt for one iteration. Default: returns goal verbatim (the classic fixed-prompt Ralph loop).
onIteration(ctx) => void | Promise<void>Observability hook fired at the end of each iteration.
abortSignalAbortSignalChecked before each iteration and propagated into every agent.run().

Returns { iterations, done, lastText, stopReason } where stopReason is 'done' | 'max_iterations' | 'aborted'. It is pure composition over the public API: zero changes to loop/session, and it does not even import createAgent (only the Agent type), so the core stays runtime-agnostic.

import { OpenAIProvider, createLocalAgent, runGoalLoop } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'

const result = await runGoalLoop({
  goal: 'Make every test in the project pass.',
  maxIterations: 6,
  // FRESH agent per iteration → clean context. Same cwd = shared state on disk.
  createIterationAgent: () => createLocalAgent({
    cwd: process.cwd(),
    provider: new OpenAIProvider({
      model: 'myModel',
      baseURL: 'http://127.0.0.1:11434/v1',
      apiKey: 'myApiKey',
    }),
    tools: createCodingTools(),
  }),
  // Cheap stop condition in TS: 0 tokens.
  isDone: () => fs.existsSync('.tests-green'),
})

console.log(result.stopReason, result.iterations)
Heartbeat vs. goal loop. These are the two flavours of the /loop other tools expose. Heartbeat = recurring by time or condition, same session accumulating context (like Claude Code /loop --interval). Goal loop = goal pursuit with a fresh context each round and state on disk (like Codex /goal).

Persistence & branching

A SessionManager persists state (messages + metadata) and lets you resume and branch conversations.

FileSessionManager node

import { FileSessionManager } from 'astorlm'

const sm = new FileSessionManager({ dir: './.astor-sessions' })

// resume
const session = await createLocalAgent({
  provider, sessionManager: sm, sessionId: 'abc-123',
})

// native simplified branching (fork)
const childSession = await session.fork({
  newSessionId: 'abc-124',
  branchFromMessageId: 'msg-15',
})

Format: an <id>.meta.json holding metadata plus an <id>.jsonl with one message per line. Auto-saved on every assistant_message, turn_end and session_end.

Also available: InMemorySessionManager (same interface, nothing on disk) for tests and ephemeral runs, and sm.list() / sm.delete(id) to enumerate and drop sessions.


Skills

Loadable instruction packs (Agent Skills spec style): Markdown blocks with YAML frontmatter (the metadata block at the top of the file, between --- lines) that the agent can consult for specific tasks.

Activation modes

ModeHow it worksWhen to use it
'filesystem'The system prompt lists name + description + path. The agent reads the body with the standard read tool.Recommended default. No meta-tools, maximum compatibility.
'on-demand'Auto-registers a load_skill meta-tool that returns the body as a tool_result.Strong models (Claude/GPT-4 class). Useful with non-filesystem sources.
'all'Concatenates every body into the system prompt at init.Few skills (≤5), always relevant.

Example from the filesystem

import { createFileSystemSkillSource } from 'astorlm'

const session = await createLocalAgent({
  provider,
  tools: createCodingTools(),                       // read is required
  skillSources: [createFileSystemSkillSource({ dir: './skills' })],
  skillMode: 'filesystem',
})

Expected layout: skills/<name>/SKILL.md with frontmatter name (kebab-case, must match the folder) and description (no XML tags).

Layered skill sources node

createLayeredSkillSource merges several skill directories into a single source with precedence, so a skill defined closer to the project overrides a same-named one defined more globally — the usual user / project / repo cascade.

import { createLayeredSkillSource } from 'astorlm'

const skills = createLayeredSkillSource({
  // ordered from LOWEST to HIGHEST precedence
  layers: ['~/.astor/skills', './.astor-skills', './skills'],
  onOverride: ({ name, winner, loser }) =>
    console.log(`skill "${name}": ${winner} shadows ${loser}`),
})

Because it presents itself as one source, the SkillRegistry never sees the per-layer duplicates — its throw-on-duplicate policy stays intact for genuinely ambiguous configs. onOverride keeps the resolution observable instead of silent.

allowed-tools (per-skill tool policy)

The Agent Skills spec lets a skill declare an allowed-tools frontmatter field: a comma-separated list of the tools it may use. The SDK stores it verbatim in metadata without interpreting it; these two helpers are the interpretation step, so you stay in control of what "the active skill" means in your harness.

import { parseAllowedTools, restrictToolsHook } from 'astorlm'

const allowed = parseAllowedTools(skill)   // string[] | null (null = no restriction declared)

const session = await createLocalAgent({
  provider,
  tools: createCodingTools(),
  hooks: allowed ? restrictToolsHook(allowed, { alwaysAllow: ['read'] }) : undefined,
})

parseAllowedTools returns null when the field is absent or empty — meaning "no restriction declared", which you should treat differently from an empty allowlist (which would deny everything). restrictToolsHook denies anything outside the list and feeds the model a mockResult explaining why (customizable via denyMessage).

Validation helpers are exported too: validateSkillName, validateSkillDescription, validateSkillSpec, SkillValidationError and SKILL_VALIDATION_LIMITS.


MCP node

Mounts an MCP server and registers its tools into the session.

import { mountMcpServer } from 'astorlm'

const mcp = await mountMcpServer({
  name: 'fs',
  transport: { type: 'stdio', command: 'npx', args: ['-y', '@modelcontextprotocol/server-filesystem', '/tmp'] },
})

for (const tool of mcp.tools) session.registerTool(tool)

// when you are done
await mcp.close()

Tools are prefixed as <name>__<tool>. Supported transports: stdio and http.

MCP Apps — interactive UI (SEP-1865)

An MCP tool can return, alongside the text for the model, a sandboxed UI component (a ui:// resource with mimeType text/html;profile=mcp-app) linked through _meta.ui.resourceUri. astorlm surfaces it on a side channel: the text goes to the model, the UI travels separately (it never pollutes the context).

const mcp = await mountMcpServer({
  name: 'docs',
  transport: { type: 'stdio', command: 'npx', args: ['tsx', 'server.ts'] },
  // fires when a tool returns _meta.ui.resourceUri
  onToolUi: (ui) => {
    console.log(ui.toolName, ui.resourceUri, ui.structuredContent)
  },
})

// read the ui:// template and render it host-side
const view = await mcp.readUiResource('ui://semantic/results')
// view.text = the component HTML · view.mimeType = 'text/html;profile=mcp-app'

onToolUi receives { toolName, resourceUri, structuredContent, content }. readUiResource(uri) / readResource(uri) read resources from the mounted server. Full example (a semantic server plus a host that renders the component): examples/35-mcp-apps-semantic.


Executors node

A swappable backend for running commands. The bash tools talk to this, never to child_process directly.

LocalExecutor

Native Node backend. It is the default when you use createLocalAgent.

DockerExecutor

Runs every command inside a container. Real sandboxing.

import { DockerExecutor } from 'astorlm'

const session = await createLocalAgent({
  provider,
  tools: createCodingTools(),
  executor: new DockerExecutor({
    image: 'node:20-alpine',
    workdir: '/workspace',
    mounts: [{ host: process.cwd(), container: '/workspace' }],
  }),
})

Custom

Implement the Executor interface (exec, spawn, getOutput, kill, dispose) and pass it in. Useful for remote backends (SSH, Firecracker, already-provisioned containers).

Executor vs. CodeRunner. An Executor runs arbitrary shell commands with the host toolchain, and isolation comes from the OS layer (Docker) or from nothing at all (LocalExecutor). For running generated source code rather than shell commands, the sibling primitive is CodeRunner — see WASM code sandbox.

Prompt compiler

compilePrompts assembles a system prompt out of independent prompt modules instead of one hand-maintained string. Each module carries a category and a priority; the compiler orders them by a layout, drops exact duplicate sentences, and runs a static analysis pass that surfaces likely contradictions and redundancies before they reach the model. Lives in astorlm/prompt (also re-exported from the barrel).

The point is that prompt fragments coming from different places — your product identity, a tenant policy, a per-task constraint — can be authored separately and still produce one coherent prompt, with the conflicts reported rather than silently merged.

import { compilePrompts, formatPromptReport } from 'astorlm/prompt'

const { systemPrompt, report, print } = compilePrompts({
  modules: [
    { id: 'identity',  category: 'identity',   content: 'You are a release assistant.' },
    { id: 'tone',      category: 'format',     content: 'Answer in short bullet points.' },
    { id: 'safety',    category: 'policy',     content: 'Never push to main.', priority: 10 },
    { id: 'repo',      category: 'context',    content: 'The repo uses pnpm workspaces.' },
  ],
  // layout: default is ['identity', 'context', 'constraint', 'policy', 'format']
  deduplicate: true,
  detectConflicts: true,
})

console.log(report.conflicts)       // [{ type, severity, description, moduleIds }]
console.log(report.tokenEstimate)  // approximate size of the compiled prompt
console.log(print())               // or formatPromptReport(report) — human-readable audit

const agent = await createLocalAgent({ provider, systemPrompt })

Options (PromptCompilerOptions)

OptionTypeNotes
modulesPromptModule[]Required. Each one: { id, content, category, name?, priority? }.
layoutstring[]Category order in the final prompt. Default ['identity', 'context', 'constraint', 'policy', 'format'] — identity at the top, critical policies near the bottom.
deduplicatebooleanRemoves identical sentences across modules. Default true.
detectConflictsbooleanStatic analysis for contradictions and high redundancy. Default true.

priority (default 0) breaks ties: on an exact duplicate id or item, the higher priority wins. The returned report is { conflicts, modulesCompiled, tokenEstimate? }, and each conflict is { type: 'contradiction' | 'redundancy', severity: 'high' | 'medium' | 'low', description, moduleIds }.

The base prompt builder is exported alongside it: buildSystemPrompt and DEFAULT_SYSTEM_PROMPT.

Context optimizer

Automatic history compaction as the context approaches the model's limit.

await createLocalAgent({
  provider, tools,
  contextOptimizer: {
    maxTokens: 128_000,        // when the model does not expose contextLimit
    compressThreshold: 0.8,    // optimize at 80% of the limit
    keepRecentTurns: 3,
  },
})

Pass false to disable it. It is structural (prune/dedupe), not LLM summarization.

Retry

Automatically retries transient model errors (429, 5xx, timeouts).

await createLocalAgent({
  provider, tools,
  retry: {
    maxAttempts: 3,
    baseDelayMs: 500,
    maxDelayMs: 5000,
    jitter: true,
  },
})

It does not retry once the stream has emitted any event (which would duplicate already-streamed text). Emits provider_retry on the bus. The classification helpers are public: isTransientError and computeBackoffDelay.

Token usage

session.getUsage()
// { inputTokens, outputTokens, cacheReadTokens?, cacheCreationTokens? }

Accumulated across every run(). Each turn_end also carries that turn's usage. No pricing here — compute cost yourself, or use metrics.


Tracing & OpenTelemetry experimental

An observability layer: it derives a tree of spans (trace segments with a start and an end) from the agent's event bus, without touching the loop. It lives under astorlm/experimental/tracing (core-agnostic, using Web Crypto for ids), and the OTLP exporter under astorlm/experimental/tracing/otel.

One trace per agent.run(), with this hierarchy:

session                          // one run (agent.run)
  └─ turn                        // each turn → tokens + finish_reason
       ├─ provider_call          // model call → TTFT (time-to-first-token)
       └─ tool_execution         // each tool → duration, is_error

API

SymbolImportWhat it does
attachTracer(agent, { exporter, now? })astorlm/experimental/tracingSubscribes to the bus and builds the spans. Returns { detach(), currentTraceId() }. 100% additive — it does not modify the loop.
createInMemoryExporter()astorlm/experimental/tracingCollects spans into an array (.spans, .root(), .childrenOf(id)). For tests and local inspection.
createTracer(opts?)astorlm/experimental/tracingLow-level span factory (shared trace id + OTel-style ids). Normally handled by attachTracer.
createOtlpSpanExporter(opts)astorlm/experimental/tracing/otelOTLP/HTTP (JSON) exporter implemented with nothing but fetch — no OpenTelemetry SDK. Sends to any OTLP collector.

Span attributes

Where it makes sense, OpenTelemetry's GenAI semantic conventions are used (gen_ai.* keys), so the OTLP exporter maps almost 1:1:

  • gen_ai.system, gen_ai.request.model — provider and model.
  • gen_ai.usage.input_tokens / output_tokens / cache_read_tokens — tokens for the turn.
  • gen_ai.response.finish_reasons — why the turn stopped.
  • astor.ttft_ms — time-to-first-token of the model call.
  • astor.tool.name / duration_ms / is_error — per tool.
  • astor.provider_retries, astor.steering_count, astor.session.end_reason.

In-memory tracing

import { attachTracer, createInMemoryExporter } from 'astorlm/experimental/tracing'

const exporter = createInMemoryExporter()
const tracer = attachTracer(agent, { exporter })   // one line, does not touch the loop

await agent.run('List the .ts files in this project.')
tracer.detach()

for (const s of exporter.spans) {
  console.log(s.kind, s.name, s.endTime! - s.startTime, 'ms', s.attributes)
}

Exporting to an OpenTelemetry collector (OTLP)

import { attachTracer } from 'astorlm/experimental/tracing'
import { createOtlpSpanExporter } from 'astorlm/experimental/tracing/otel'

const exporter = createOtlpSpanExporter({
  endpoint: 'http://localhost:4318/v1/traces',  // Collector / Tempo / Jaeger / Honeycomb / Phoenix / Langfuse
  serviceName: 'my-agent',
  headers: { Authorization: 'Bearer ...' },          // optional (vendor keys)
  onError: (err) => console.warn('export failed', err),
})

const tracer = attachTracer(agent, { exporter })
await agent.run('...')
await exporter.shutdown()   // final flush before exiting
No heavy dependencies. The OTLP exporter serializes to OTLP/HTTP JSON using native fetch (Node ≥ 20 / browser). It buffers spans and flushes by batch size (maxBatch, default 256) or by timer (flushIntervalMs, default 5s, with unref() so it does not keep the event loop alive). Hex ids are passed through as-is (the OTLP/JSON convention for trace_id/span_id).

Runnable examples: examples/24-tracing-basic (in-memory tree + dashboard, pnpm start 24) and examples/25-tracing-otel (export to a collector, pnpm start 25).


Metrics & cost experimental

Aggregates operational metrics for the run off the event bus, and optionally computes cost in USD. Lives under astorlm/experimental/metrics (core-agnostic). No hardcoded prices — you supply the table.

API

SymbolWhat it does
attachMetrics(agent, { pricing?, now? })Subscribes to the bus and aggregates counters, tokens, cost and latencies. Returns { snapshot(), reset(), detach() }.
computeCost(usage, pricing)USD for a TokenUsage under a ModelPricing (prices per million tokens).
resolvePricing(table, model)Looks up the model's pricing: exact match first, then longest prefix (gpt-4o matches gpt-4o-2024-...).

The snapshot() includes: runs, turns, providerCalls, toolCalls, toolErrors, providerRetries, tokens (in/out/cache), costUsd? (when pricing is supplied) and latency with stats (count/sum/min/max/avg) for TTFT, per-turn duration and per-tool duration.

import { attachMetrics } from 'astorlm/experimental/metrics'

const metrics = attachMetrics(agent, {
  pricing: {                                  // USD per 1,000,000 tokens
    myModel: { inputPer1M: 0.5, outputPer1M: 1.5 },
    'gpt-4o': { inputPer1M: 2.5, outputPer1M: 10, cacheReadPer1M: 1.25 },
  },
})

await agent.run('...')
const m = metrics.snapshot()
console.log(m.costUsd, m.latency.ttftMs.avg, m.toolErrors)

Example: examples/26-metrics-cost (pnpm start 26).


Record & Replay experimental

Records exactly the events the model produced and plays them back later with no network and no tokens — the deterministic debugging primitive. Lives under astorlm/experimental/replay (core-agnostic; a Recording is a serializable object, so persistence is JSON.stringify plus your storage).

API

SymbolWhat it does
createRecordingProvider(inner, opts?)Wraps a real Provider: passes events through to the loop while capturing them turn by turn. getRecording() returns a serializable object.
createReplayProvider(recording, opts?)A Provider that replays the recording turn by turn. onExhausted: 'throw' | 'end' controls what happens if the loop asks for more turns than were recorded.
import { createRecordingProvider, createReplayProvider } from 'astorlm/experimental/replay'

// record
const rec = createRecordingProvider(realProvider)
const agent = await createLocalAgent({ provider: rec, tools })
await agent.run('...')
fs.writeFileSync('run.json', JSON.stringify(rec.getRecording()))

// replay later — same events, no model
const recording = JSON.parse(fs.readFileSync('run.json', 'utf8'))
const replay = await createLocalAgent({ provider: createReplayProvider(recording), tools })
await replay.run('...')

Capture happens at the provider boundary, so it records exactly what the model emitted, independent of tools, hooks and timing. Example: examples/27-replay (pnpm start 27).


Evals experimental

An offline evaluation harness: it runs a dataset of cases against fresh agents, applies scorers (graders), and aggregates a report with an overall pass rate and a per-scorer breakdown. Ideal for CI gating. Lives under astorlm/experimental/evals.

Bundled scorers

ScorerWhat it measures
exactMatch(opts?)Equality with case.expected (normalizes case and whitespace by default).
contains(substr?, opts?)The output contains the expected substring.
regexMatch(re)The output matches the regular expression.
toolTrajectory(tools, { mode })The sequence of tools called. Modes: 'exact', 'ordered-subset', 'set'.
llmJudge({ provider, rubric, passThreshold? })A model scores the answer against a rubric (0..1). Robust to noisy output.
import { runEval, contains, toolTrajectory, llmJudge } from 'astorlm/experimental/evals'

const report = await runEval({
  dataset: [{ id: 'q1', input: 'How many .ts files are there?', expected: 'a number' }],
  concurrency: 1,
  createAgent: (c) => createLocalAgent({ provider: provider(), tools: createReadOnlyTools() }),
  scorers: [
    contains('.ts'),
    toolTrajectory(['ls'], { mode: 'set' }),
    llmJudge({ provider: provider(), rubric: 'Pass if it answers with real data.' }),
  ],
})

console.log(report.summary.passRate, report.summary.byScorer)
// CI gate: if (report.summary.passRate < 0.8) process.exit(1)
Every case runs in a fresh agent. createAgent(case) is invoked per case so they never share history. Combine it with createReplayProvider for fast, network-free regression runs. llmJudge takes a Provider — point it at your own OpenAI-compatible endpoint.

Example: examples/28-evals (pnpm start 28).


WASM code sandbox experimental

CodeRunner is the sibling primitive to Executor: instead of running shell commands, it evaluates a self-contained source snippet inside a memory-safe WebAssembly sandbox. Isolation is capability-based — the guest has no filesystem, no network and no host syscalls unless explicitly granted. Because the engine is pure WASM it runs anywhere WASM does (Node, Deno, browser, edge), which is exactly where DockerExecutor cannot reach (no daemon, no child_process).

The interfaces (CodeRunner, RunCodeOptions, RunCodeResult) live in the agnostic core; the concrete implementation lives under astorlm/experimental/wasm-runner, backed by QuickJS compiled to WebAssembly.

import { QuickJsCodeRunner, createCodeRunnerTool } from 'astorlm/experimental/wasm-runner'

const runner = new QuickJsCodeRunner({
  timeoutMs: 5000,              // wall-clock budget per run
  memoryLimitBytes: 64 * 1024 * 1024,
  maxOutputBytes: 200_000,
})

// 1) direct use
const res = await runner.run({
  code: 'const xs = [1,2,3]; console.log(xs.map(x => x * 2)); xs.length',
  globals: { threshold: 10 },   // JSON-serializable values copied read-only into the guest
})
res.ok, res.value, res.stdout, res.timedOut, res.outOfMemory, res.durationMs

// 2) as an agent tool — the model writes code, it runs with no host access
const agent = await createLocalAgent({
  provider,
  tools: [createCodeRunnerTool({ runner })],   // tool name defaults to run_code
})

await runner.dispose()

RunCodeOptions

FieldTypeNotes
codestringThe source to evaluate. Required.
timeoutMsnumberWall-clock budget enforced by an interrupt handler. Default 5_000.
memoryLimitBytesnumberHard memory cap for the guest runtime. Default 64 MB.
maxOutputBytesnumberMax captured stdout/stderr before truncation. Default 200 KB.
globalsRecord<string, unknown>JSON-serializable values copied read-only into the guest scope. Host functions are intentionally not supported, to keep the isolation guarantee airtight.
abortSignalAbortSignalCancels the run.

RunCodeResult

{ ok, value?, stdout, stderr, error?, truncated, timedOut, outOfMemory, durationMs } — ok is true only if the code ran to completion without throwing, timing out or running out of memory, and the three failure flags tell the three cases apart.

Opt-in, and not a replacement for Executor. createCodeRunnerTool is not part of createCodingTools(); register it explicitly when you want the agent to evaluate code it generated without touching the host. Each run() instantiates and disposes its own runtime and context, so nothing leaks between calls; the WASM module itself is loaded once and cached (lazily, so importing the module does not pull the binary until a run actually happens). It runs snippets, not shell commands — for npm, git or python you still want an Executor.

Example: examples/32-wasm-runner (pnpm start 32). Requires the quickjs-emscripten peer dependency.


Agent contract experimental

Establishes deterministic consumption limits (budget) and safety guardrails (sandbox) at runtime. Validation happens through hooks, before calling the LLM or executing any tool.

import { createAgent } from 'astorlm'
import { createContractHooks } from 'astorlm/experimental/contract'

const contract = {
  budget: {
    maxTurns: 5,
    maxTotalTokens: 30000,
  },
  sandbox: {
    allowedPaths: ['src/**/*'],
    bash: {
      deniedCommands: ['rm -rf'],
    }
  }
}

const agent = await createAgent({
  provider,
  hooks: createContractHooks(contract) // <-- hooks middleware
})

Contract options

SectionOptionTypeDescription
budgetmaxTurnsnumberCap on total turns allowed in the session.
maxInputTokensnumberCumulative input token limit.
maxOutputTokensnumberCumulative output token limit.
maxTotalTokensnumberLimit on input + output tokens combined.
maxDurationMsnumberMaximum session execution time, in milliseconds.
toolsallowstring[]Allowed tools (whitelist).
denystring[]Explicitly blocked tools.
sandboxallowedPathsstring[]Allowed filesystem glob patterns (e.g. ['src/**/*']).
deniedPathsstring[]Blocked filesystem glob patterns.
bash{ allowedCommands?, deniedCommands? }Approved shell commands (exact match or prefix) and denied keywords (as strings, globs or regexes).

If the session exceeds its budget or a tool violates the sandbox, the hooks throw a ContractViolationError and emit a contract_violation event on the session's event bus.


Error registry experimental

A federated store of errors and approved resolutions. When an agent hits an error that already has a known fix, it receives it as a hint on the next tool_result.

import { createErrorRegistry, errorRegistryHooks } from 'astorlm/experimental/error-registry'

const registry = await createErrorRegistry({
  storePath: './.astor-errors.jsonl',
})

const session = await createLocalAgent({
  provider, tools,
  hooks: errorRegistryHooks({ registry }),
})

// human-in-the-loop: approve a candidate resolution
await registry.approveResolution('res-123', 'reviewer-name')

Matching is done by cosine similarity over embeddings (OpenAI-compatible) with a fuzzy Jaccard fallback. Paths, UUIDs, IPs and timestamps are normalized before comparison.

Unstable API. It lives under astorlm/experimental/* on purpose — it may change between minor versions.

edge-boost (weak models) experimental

edgeBoost(options, tuning?) is an options transformer (a pure function CreateAgentOptions → CreateAgentOptions) that hardens an agent for weak models: OpenRouter free tiers, small local models (Ollama/vLLM), heavily quantized checkpoints. It defends against the failure mode where the model, in a long multi-turn context with tools (typically the synthesis turn after several rounds), degenerates into a repetition loop or hangs without producing content until it hits a timeout.

import { createLocalAgent } from 'astorlm/core'
import { OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
import { edgeBoost } from 'astorlm/experimental/edge-boost'

// Wrap the same options you already pass to createLocalAgent. One line, no extra layers.
const agent = await createLocalAgent(edgeBoost({
  provider: new OpenAIProvider({
    model: 'myModel',
    baseURL: 'http://127.0.0.1:11434/v1',
    apiKey: 'myApiKey',
  }),
  tools: createCodingTools(),
}))

// Tune any section; pass `false` to turn one off. E.g. force synthesis earlier.
const tuned = edgeBoost(options, { synthesis: { forceAfterToolRounds: 2 } })

Four opt-in defenses; none of them changes default behaviour unless you enable it:

  • Degeneration guard (createGuardedProvider) — detects repeated identical deltas, content-free stalls and reasoning-budget blowout; aborts cheaply and retries against the same endpoint. It owns retries caused by degeneration; network/HTTP retries still belong to the loop's retry. Failover between models is deliberately out of scope (that belongs to the routing proxy — see Providers).
  • Context diet — truncates old tool_results per call (idempotent, and compatible with the core optimizer's markers) plus an optional compact system prompt; it never mutates history. The optional digest summarizes an evicted block with an injected cheap provider instead of truncating it (cached per block, falling back to truncation on failure or timeout).
  • Forced synthesis — after N tool rounds in the current request, it appends an instruction and forces toolChoice: 'none' (or strips the tools) so the model stops calling tools and answers.
  • Sampling defaults — injects max_tokens and frequencyPenalty (repetition penalty) when the call does not carry them. Plus optional semantic tool pruning (top-K relevant tools per turn via an embedder; disabled by default, degrading gracefully when no embedder is present).

Exports: edgeBoost, edgeBoostHooks, createGuardedProvider, DegenerationError, mergeSessionHooks (hook composition, useful in general), buildContextDietHook/buildSynthesisHook/buildToolPruningHook, resolveEdgeBoostTuning plus the tuning types. Runnable example: 34-edge-boost (pnpm start 34, with optional --prune / --digest flags).

Unstable API. It lives under astorlm/experimental/* on purpose — it is not re-exported from the main barrel and may change between minor versions.