Deterministic testing for LLM agents and non-deterministic pipelines.
About • Features • The Pipeline • Installation • Quickstart • Examples & Output
Tardi is an open-source testing framework engineered specifically for evaluating agentic workflows, autonomous LLM scripts, and non-deterministic applications.
Testing AI agents presents a unique challenge: traditional testing frameworks cannot evaluate non-deterministic natural language outputs, and raw "LLM-as-a-judge" evaluation pipelines are prohibitively expensive and prone to hallucination at scale.
Tardi solves this by implementing a tiered assertion gauntlet. Instead of blindly sending every agent execution trace to an evaluation model, Tardi enforces strict deterministic constraints first. If your agent crashes, hangs in an infinite loop, or returns malformed JSON, Tardi fails the test immediately—preventing unnecessary LLM API calls and accelerating your feedback loop.
- Tiered Evaluation Engine: Catch process crashes, timeouts, and schema mismatches deterministically before triggering expensive LLM judges.
- Deterministic Prompt Assembly & Token Budgeting: Automatically enforce token limits when assembling prompts with context, constraints, and dynamic inputs. Avoid context overflow dynamically.
- Injection Defense Layer: Built-in semantic and tag-based sanitization automatically neutralizes prompt injection attempts (e.g., users attempting to escape
<user_input>blocks). - Concurrency & Rate-Limiting: Execute test suites in parallel with intelligent chunking to maximize throughput without exceeding API rate limits.
- Provider Agnostic: Built on the standard Vercel AI SDK, Tardi supports hot-swappable evaluation models from OpenAI, Google, Anthropic, and local endpoints.
- Interactive REPL & Natural Language CLI: Includes a CLI environment for rapid test synthesis, execution, and debugging, powered by an onboard NLP intent parser.
- Secure Credential Management: Safely stores provider API keys in your native OS keychain during local development.
When you execute an iteration, Tardi evaluates the agent's output through a strict, cost-saving pipeline. Deterministic checks always run first.
graph TD
Start([Run Agent]) --> ProcessCheck{Process Crashed?}
ProcessCheck -- Yes --> FailCrash[FAIL: CRASH]
ProcessCheck -- No --> TimeoutCheck{Exceeded Timeout?}
TimeoutCheck -- Yes --> FailTimeout[FAIL: TIMEOUT]
TimeoutCheck -- No --> RegexCheck{Matches Regex?}
RegexCheck -- No --> FailRegex[FAIL: REGEX_MISMATCH]
RegexCheck -- Yes --> SchemaCheck{Valid JSON Schema?}
SchemaCheck -- No --> FailSchema[FAIL: SCHEMA_MISMATCH]
SchemaCheck -- Yes --> LLMJudge{LLM Judge Approves?}
LLMJudge -- No --> FailJudge[FAIL: LLM_JUDGE_FAIL]
LLMJudge -- Yes --> Pass([PASS])
style FailCrash fill:#fbd4d4,stroke:#f87171,color:#000
style FailTimeout fill:#fbd4d4,stroke:#f87171,color:#000
style FailRegex fill:#fbd4d4,stroke:#f87171,color:#000
style FailSchema fill:#fbd4d4,stroke:#f87171,color:#000
style FailJudge fill:#fbd4d4,stroke:#f87171,color:#000
style Pass fill:#dcfce7,stroke:#4ade80,color:#000
Install the CLI globally via npm to use the tardi command anywhere:
npm install -g @abdur-raheem/tardiTardi can automatically clone an agent repository, detect its entry point, and synthesize a test gauntlet based on golden runs.
tardi github https://github.com/Nyx-abu/demo-agent.gitIf you prefer to configure your suite manually:
tardi init
tardi auth login google
tardi run tests/Tardi utilizes simple YAML files (*.tardi.yaml) to define test suites. Below are detailed examples of how Tardi behaves under different failure and success conditions.
A common requirement for AI agents is to return structured JSON. Tardi catches malformed JSON or missing fields deterministically, bypassing the LLM Judge entirely.
Test Configuration:
# tests/json-schema.tardi.yaml
name: Agent JSON Output Test
command:
executable: node
args: ["src/agent.js"]
iterations: 5
concurrency: 2
timeoutMs: 30000
assertions:
jsonSchema:
type: object
required: ["status", "result"]Expected Tardi Output (Schema Mismatch):
If the agent forgets to include the result field, Tardi immediately aborts the pipeline and outputs a schema mismatch error:
╭──────────────────────── Tardi Execution Summary ────────────────────────╮
│ ┌─────────────┬──────────┐ │
│ │ Metric │ Value │ │
│ ├─────────────┼──────────┤ │
│ │ Total Runs │ 5 │ │
│ │ Passed │ 4 │ │
│ │ Failed │ 1 │ │
│ │ Flakiness │ 20.00% │ │
│ └─────────────┴──────────┘ │
│ │
│ Failures: │
│ - Iteration 3 [SCHEMA_MISMATCH]: Output failed JSON Schema validation. │
│ Missing required property: 'result'. │
╰───────────────────────────────────────────────────────────────────────────╯
If your agent passes all deterministic constraints, Tardi forwards the raw output to an LLM evaluator. The LLM evaluator uses your custom rubric to score the output.
Test Configuration:
# tests/evaluator.tardi.yaml
name: Semantic Agent Test
command:
executable: node
args: ["src/summarize.js"]
iterations: 3
assertions:
regex: "summary:"
evaluator:
provider: google
model: gemini-2.5-flash
rubric: |
Evaluate if the agent successfully summarized the input document.
Return only 'PASS' or 'FAIL'.Expected Tardi Output (LLM Judge Failure):
When testing complex autonomous scripts, agents often crash mid-execution. Tardi intercepts process crashes and records the exit codes, allowing you to debug stability issues over multiple iterations.
Expected Tardi Output (Crash):
╭──────────────────────── Tardi Execution Summary ────────────────────────╮
│ Failures: │
│ - Iteration 2 [CRASH]: Agent process exited with code 1. │
│ Stderr: Error: Uncaught exception in src/agent.js:42 │
╰───────────────────────────────────────────────────────────────────────────╯
Tardi is built from the ground up for seamless enterprise integration. Whether deploying in a regulated environment or a hyper-scale pipeline, Tardi ensures your tests remain secure and performant.
- Deterministic assertions run locally.
- Evaluator calls may send selected agent output to the configured LLM provider.
- API keys are resolved from environment variables or the OS keychain.
- CI users should inject provider keys through their CI secret store.
- Enterprise secret-manager, SSO, and tamper-proof audit-log integrations are not implemented yet.
Inject API keys securely via your CI runner environment variables to execute your gauntlets:
GOOGLE_GENERATIVE_AI_API_KEY="your-key-here" tardi run tests/Pull requests are welcome. For major architectural changes, please open an issue first to discuss your proposed modifications.
Please see our Contributing Guide and our Code of Conduct.
This project is licensed under the ISC License.