# AgentRun AgentRun (stylized `agent.run()`, officially "The Agentrun Workflow [[Domain Specific Languages (DSLs)|DSL]]") is an open source workflow language for [[AI Agents]] by Parcha Labs. A workflow is a document of typed steps (tool calls, code, decisions, agent calls, escalation) that an interpreter runs. Your application keeps its own tools, model access, permissions and budgets; AgentRun decides what runs next. The split is by cost. Small yes/no decisions ("does this answer fit the request?") go to [[Jev]], the [[Decision Models (DMs)|decision model]] by [[TypeSafe AI]]. The expensive agent only gets called when something needs investigating. Code reads Jev's answer and its confidence, then picks the next step. ## Why a workflow language The usual agent setup is one loop: give the model tools, then read the transcript afterwards to see what it did. Fine for exploration. For a repeatable process like answering support tickets, screening evidence or triaging requests, I want the steps written down where I can test them. In AgentRun those steps live in a workflow document you can inspect, rerun, test, or hand to an agent as a tool. The README admits that for a fixed sequence, plain functions may be enough. It starts to pay off when deterministic steps, semantic decisions and agent work sit in the same flow. ## How it works The support example from the README: 1. Search the help center for an answer (a `call` node delegated to a host tool, with a required deadline) 2. Check it with Jev (a `judge` node): the answer must have text and a source, and Jev must answer `yes` with a confidence of at least 0.8 3. If that fails, allow ONE agent attempt to investigate, then check again 4. If it still fails, escalate for review (by a human or the host application) On the scripted cases, a password reset or an invoice lookup costs zero agent calls and one decision call. A failed payment costs one agent call and two decision calls. That's [[Confidence-Gated Routing]]: Jev settles the easy tickets, and the agent and a [[Human-in-the-Loop|human]] only see what's left. The primitives follow Jev's question types ([[System One Primitives]]): - Control flow: `chain`, `parallel`, `map` (bounded concurrency) and `loop`, which requires a maximum iteration count, so no runaway loops - Semantic decisions: `judge` (a flat set of typed questions), `pick` (one item, or none), `sift` (filter a collection) and `route` (pick a branch by meaning, with an explicit "uncertain" fallback) - Generative work: `agent`, `decide` and `extract` return a schema-checked result from your agent. A generative node can declare a `verify` step that judges the candidate; a failed check is never silently accepted - Plumbing: `code` (trusted JavaScript), `call` (tools, executors, shell), `workflow` (child workflows with their own input/output contracts), `escalate`, `report` and `artifact` Workflows are written as JSON or in [[TypeScript]] with a builder and Zod contracts (the builder infers input and output types). A CLI validates and dry-runs them, and a [[Pi Mono|Pi]] extension lets the coding agent build, inspect and run workflows from a prompt (`/agentrun ...`). ## Status - Created on 23 September 2026, about a week after TypeSafe launched Jev; around 90 GitHub stars at the time of writing, actively developed - Beta (`0.1.0-beta.4`): `npm install @parcha/agentrun-dsl@beta`, Node 22.19+; the API may still change - Three packages: the DSL core, a Jev adapter and the Pi extension - Apache-2.0, copyright Parcha Labs, "built by Grep.ai" - Live Jev calls need a TypeSafe API key; workflows without Jev decisions don't ## Limits From the docs: - `code` nodes execute JavaScript with your process's privileges, and validation can run probes. If you don't trust whoever wrote the workflow, the host has to sandbox it - Typed decisions and validated output shapes do NOT prove an answer is factually correct. Verification questions only judge the evidence they're given - Tools, permissions, delivery, durable storage, budgets and recovery are host hooks; AgentRun provides none of them. There's no exactly-once delivery either: an effect still pending at the deadline is reported as uncertain and never retried automatically - Calibration is your job. The 0.8 threshold in the example is the author's pick, and Anthony Maio's point about Jev applies here: individually [[AI Model Calibration|calibrated]] judgments don't add up to a calibrated workflow once you chain them through thresholds and branches. Test thresholds on labeled cases from your own task ([[Shadow Evaluation]]) ## My take TypeSafe's manifesto talks about "programming languages built on top of Jev". AgentRun is one of the first I've seen, and it comes from a third party, not from TypeSafe. It uses Jev as the cheap `if` between deterministic code and an expensive agent, which is where I think a decision model belongs. It also leaves your agent and your [[AI Agent Harness|harness]] as they are and puts bounds around them: deadlines, iteration limits, explicit escalation. That's how I see [[Agentic Engineering]] going. Whatever is repeatable moves out of the prompt and into code, and the model keeps the judgment calls. I'd start small, though. It's a beta, the decision layer is one closed API, and you build the host integration yourself. Pick one well-understood, low-stakes process (triage, routing, evidence screening) before trusting it with anything consequential. ## References - [AgentRun on GitHub (Parcha-ai/agentrun)](https://github.com/Parcha-ai/agentrun) - [Build with agent.run() (AgentRun guide)](https://github.com/Parcha-ai/agentrun/blob/main/docs/guide.md) - [@parcha/agentrun-dsl on npm](https://www.npmjs.com/package/@parcha/agentrun-dsl) - [TypeSafe documentation](https://docs.typesafe.ai/) ## Related - [[Jev]] - [[TypeSafe AI]] - [[System One Primitives]] - [[Decision Models (DMs)]] - [[Confidence-Gated Routing]] - [[Model routing]] - [[AI Guardrails]] - [[LLM Structured Outputs]] - [[Shadow Evaluation]] - [[Human-in-the-Loop]] - [[AI Agents]] - [[AI Agent Harness]] - [[Agentic loops]] - [[Agentic Engineering]] - [[Pi Mono]] - [[TypeScript]] - [[Use code for control flow, agents for reasoning]] - [[AI Model Cascades]] - [[Atomic Question Decomposition]] - [[DSLs Make LLM Output Reliable]] - [[Distinction between AI Agents and Automation Workflows]] - [[Mastra Workflows]]