# AgentRun
AgentRun (stylized `agent.run()`, officially "The Agentrun Workflow [[Domain Specific Languages (DSLs)|DSL]]") is an open source workflow language for [[AI Agents]] by Parcha Labs. A workflow is a document of typed steps (tool calls, code, decisions, agent calls, escalation) that an interpreter runs. Your application keeps its own tools, model access, permissions and budgets; AgentRun decides what runs next.
The split is by cost. Small yes/no decisions ("does this answer fit the request?") go to [[Jev]], the [[Decision Models (DMs)|decision model]] by [[TypeSafe AI]]. The expensive agent only gets called when something needs investigating. Code reads Jev's answer and its confidence, then picks the next step.
## Why a workflow language
The usual agent setup is one loop: give the model tools, then read the transcript afterwards to see what it did. Fine for exploration. For a repeatable process like answering support tickets, screening evidence or triaging requests, I want the steps written down where I can test them.
In AgentRun those steps live in a workflow document you can inspect, rerun, test, or hand to an agent as a tool. The README admits that for a fixed sequence, plain functions may be enough. It starts to pay off when deterministic steps, semantic decisions and agent work sit in the same flow.
## How it works
The support example from the README:
1. Search the help center for an answer (a `call` node delegated to a host tool, with a required deadline)
2. Check it with Jev (a `judge` node): the answer must have text and a source, and Jev must answer `yes` with a confidence of at least 0.8
3. If that fails, allow ONE agent attempt to investigate, then check again
4. If it still fails, escalate for review (by a human or the host application)
On the scripted cases, a password reset or an invoice lookup costs zero agent calls and one decision call. A failed payment costs one agent call and two decision calls. That's [[Confidence-Gated Routing]]: Jev settles the easy tickets, and the agent and a [[Human-in-the-Loop|human]] only see what's left.
The primitives follow Jev's question types ([[System One Primitives]]):
- Control flow: `chain`, `parallel`, `map` (bounded concurrency) and `loop`, which requires a maximum iteration count, so no runaway loops
- Semantic decisions: `judge` (a flat set of typed questions), `pick` (one item, or none), `sift` (filter a collection) and `route` (pick a branch by meaning, with an explicit "uncertain" fallback)
- Generative work: `agent`, `decide` and `extract` return a schema-checked result from your agent. A generative node can declare a `verify` step that judges the candidate; a failed check is never silently accepted
- Plumbing: `code` (trusted JavaScript), `call` (tools, executors, shell), `workflow` (child workflows with their own input/output contracts), `escalate`, `report` and `artifact`
Workflows are written as JSON or in [[TypeScript]] with a builder and Zod contracts (the builder infers input and output types). A CLI validates and dry-runs them, and a [[Pi Mono|Pi]] extension lets the coding agent build, inspect and run workflows from a prompt (`/agentrun ...`).
## Status
- Created on 23 September 2026, about a week after TypeSafe launched Jev; around 90 GitHub stars at the time of writing, actively developed
- Beta (`0.1.0-beta.4`): `npm install @parcha/agentrun-dsl@beta`, Node 22.19+; the API may still change
- Three packages: the DSL core, a Jev adapter and the Pi extension
- Apache-2.0, copyright Parcha Labs, "built by Grep.ai"
- Live Jev calls need a TypeSafe API key; workflows without Jev decisions don't
## Limits
From the docs:
- `code` nodes execute JavaScript with your process's privileges, and validation can run probes. If you don't trust whoever wrote the workflow, the host has to sandbox it
- Typed decisions and validated output shapes do NOT prove an answer is factually correct. Verification questions only judge the evidence they're given
- Tools, permissions, delivery, durable storage, budgets and recovery are host hooks; AgentRun provides none of them. There's no exactly-once delivery either: an effect still pending at the deadline is reported as uncertain and never retried automatically
- Calibration is your job. The 0.8 threshold in the example is the author's pick, and Anthony Maio's point about Jev applies here: individually [[AI Model Calibration|calibrated]] judgments don't add up to a calibrated workflow once you chain them through thresholds and branches. Test thresholds on labeled cases from your own task ([[Shadow Evaluation]])
## My take
TypeSafe's manifesto talks about "programming languages built on top of Jev". AgentRun is one of the first I've seen, and it comes from a third party, not from TypeSafe. It uses Jev as the cheap `if` between deterministic code and an expensive agent, which is where I think a decision model belongs.
It also leaves your agent and your [[AI Agent Harness|harness]] as they are and puts bounds around them: deadlines, iteration limits, explicit escalation. That's how I see [[Agentic Engineering]] going. Whatever is repeatable moves out of the prompt and into code, and the model keeps the judgment calls.
I'd start small, though. It's a beta, the decision layer is one closed API, and you build the host integration yourself. Pick one well-understood, low-stakes process (triage, routing, evidence screening) before trusting it with anything consequential.
## References
- [AgentRun on GitHub (Parcha-ai/agentrun)](https://github.com/Parcha-ai/agentrun)
- [Build with agent.run() (AgentRun guide)](https://github.com/Parcha-ai/agentrun/blob/main/docs/guide.md)
- [@parcha/agentrun-dsl on npm](https://www.npmjs.com/package/@parcha/agentrun-dsl)
- [TypeSafe documentation](https://docs.typesafe.ai/)
## Related
- [[Jev]]
- [[TypeSafe AI]]
- [[System One Primitives]]
- [[Decision Models (DMs)]]
- [[Confidence-Gated Routing]]
- [[Model routing]]
- [[AI Guardrails]]
- [[LLM Structured Outputs]]
- [[Shadow Evaluation]]
- [[Human-in-the-Loop]]
- [[AI Agents]]
- [[AI Agent Harness]]
- [[Agentic loops]]
- [[Agentic Engineering]]
- [[Pi Mono]]
- [[TypeScript]]
- [[Use code for control flow, agents for reasoning]]
- [[AI Model Cascades]]
- [[Atomic Question Decomposition]]
- [[DSLs Make LLM Output Reliable]]
- [[Distinction between AI Agents and Automation Workflows]]
- [[Mastra Workflows]]