Skip to content

Latest commit

 

History

177 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Freyja

crates.io docs.rs CI license

A provider-neutral LLM client for Rust, and the foundation for building agents on top of it.

Warning

Under active development and not ready for production use. The public API is unstable and will change without notice before 1.0.0.

You write one request. Freyja translates it into whatever wire format the model you picked actually speaks, sends it, and translates the answer back. Changing vendor is changing one line.

let client = Client::from_env(EndpointPreset::OpenAi).expect("OPENAI_API_KEY");
// or Anthropic, or Gemini, or any compatible endpoint. Nothing else changes.

"Any compatible endpoint" includes ones whose URL follows neither the dialect nor the vendor, and gateways that want a credential of their own beside the key:

let config = EndpointConfig::new(Dialect::OpenAiChat, "acme-gw", "https://gw.acme.test/v1")
    .api_key_env("ACME_API_KEY")
    .path("/openai/deployments/gpt4/chat/completions") // replaces the dialect's path
    .query("api-version", "2024-02-01")                // pinned on every request
    .header("x-acme-tenant", "engineering")            // printed, it is configuration
    .secret_header("x-acme-passport", &passport);      // withheld, it is a credential

Freyja does the URL joining and escaping, so a request never grows a second ?. A classified value is withheld from Debug, from error messages and from config.url(). An unclassified one is withheld only if its name happens to look like a credential. For an endpoint that wants the key in the URL rather than a header, Auth::Query("key") keeps it out of config.url(). See Custom endpoints.

That matters because every vendor invented a different shape for the same ideas. A tool call is a flat item on OpenAI, a typed step on Gemini, a nested block on Anthropic, and a fourth arrangement on the Chat Completions format most other vendors copy. Your code sees none of it.

Quick start

cargo add freyja
cargo add tokio --features macros,rt-multi-thread
use freyja::{Client, GenerateRequest, Message, EndpointPreset, Role};

#[tokio::main]
async fn main() {
    let client = Client::from_env(EndpointPreset::OpenAi).expect("OPENAI_API_KEY");

    let request = GenerateRequest::new()
        .message(Message::text(Role::User, "Name three Rust crates."));

    match client.generate(&request).await {
        Ok(response) => println!("{}", response.output_text()),
        Err(error) => eprintln!("request failed: {error}"),
    }
}

Or take the same answer as it arrives. Tool-call arguments are assembled for you, so nothing hands you half a JSON object:

use freyja::StreamEvent;

let mut stream = client.stream(&request).await?;
while let Some(event) = stream.next().await? {
    match event {
        StreamEvent::TextDelta(text) => print!("{text}"),
        StreamEvent::ToolCall { name, arguments, .. } => println!("\n{name}({arguments})"),
        _ => {}
    }
}

A drained stream converts back with stream.into_response()?, so a streaming tool loop reuses the same to_message() the non-streaming one does. See Streaming.

Add typed tools and a bounded loop and you have an agent. #[tool] derives the argument schema and JSON dispatcher from an ordinary Rust function. See Building an agent.

cargo run --example simple           # one question, one answer
cargo run --example streaming        # the same answer, printed as it arrives
cargo run --example tool_loop        # a bounded agent loop
cargo run --example custom_endpoint  # an endpoint with no preset
cargo run --example retry            # a retry loop over the error classification
cargo run --example chat             # an interactive multi-turn conversation
cargo run --example portable         # one request, every vendor, and its limits
cargo run --example structured_output # JSON constrained by a schema, deserialized
cargo run --example images           # an image in a prompt, by URL or data URI
cargo run --example async_tools      # several tool calls running at once
cargo run --example agent            # the loop driven by Agent
cargo run --example guarded_tools    # tool state, run context, failures, and a guard
cargo run --example memory           # bounding what reaches the model, transcript kept whole

Documentation

Full docs in docs/, written to be read in order:

Page What it covers
Introduction What Freyja is, what it is not, and why it exists
Features What works today, and what does not
Getting started Install, set a key, make a call
Concepts The five ideas everything else follows from
Building an agent Tools, the loop, and what will bite you

Then providers, the API reference, and internals for working on Freyja itself.

Status

Phases 0 through 2 are complete, and Phase 3 has started: the neutral core is stable, four wire dialects are implemented, tool calling works end to end, typed #[tool] functions derive their schemas and dispatchers and may be sync or async, every dialect streams, failures are classified by cause, and a Storage backend decides what part of a transcript reaches the model on each turn, inside its own load.

Area State
Built-in providers OpenAI, Gemini, Anthropic, all verified against live APIs
Other endpoints DeepSeek, Groq, OpenRouter, Ollama and friends via Client::custom
Non-conventional URLs path and query reach a gateway whose URL follows neither the dialect nor the vendor
Credential safety The key is redacted wherever Freyja prints itself and marked sensitive on the wire, and secret_header and secret_query extend both to a second credential
Tool calling Typed #[tool] declarations and the full round trip
Streaming All four dialects, text verified live; tool calls offline only
Dependencies reqwest, serde, serde_json, schemars, and the companion macro crate
Errors Classified by cause, with is_retryable() and Retry-After
Untrusted endpoints Same-origin redirects only, and ceilings on body, stream, frame, retry delay and tool fan-out
Group-aware trimming window_by_groups, the rule InMemoryStorage::window uses, public for a backend of your own
Pre-flight checks client.check(&request), no network call
Structured output strict_schema() plus generate_as::<T>()
Vendor-only fields extra_for(), without forking

The workspace test suite covers the core, macro expansion, public typed-tool behavior, examples, and doctests. Features has the honest boundary, including which capabilities each provider refuses.

Roadmap

The goal: everything you need to build an AI agent in Rust, with no vendor lock-in.

Phase 0, stabilize the core. Complete. Portable defaults, tool round trips, opaque reasoning state, pooled HTTP, live verification on every provider.

Phase 1, production-grade provider layer. Complete. Four dialects, the dialect/endpoint split, streaming, typed errors, pre-flight checking, typed responses, and strict-mode schema rewriting.

Phase 2, the agent. Complete. Tool and #[tool] derive schemas from sync or async function signatures and provide typed execution, and Agent drives the tool-calling loop automatically, dispatching parallel tool calls concurrently, eight at a time. Tool is now a trait, so a tool can hold state in its fields, be built at runtime, and report failure as text the model recovers from; Context carries per-run data to every call without exposing it to the model. Agent::guard vets every requested call before dispatch, so a policy can refuse one and the model reads why.

Phase 3, memory and context. Started. Storage is the backend a conversation reads and writes, and it decides what reaches the model by trimming inside its own load, with the caller's transcript kept whole. InMemoryStorage::window bounds one by turn group, and window_by_groups is public, so a backend of your own applies the same group-aware rule rather than reimplementing it or cutting mid-pair. Token-aware windows, summarization, persistent backends, and retrieval with embeddings and a vector store are not built.

Phase 4, orchestration. The namesake. Multi-agent handoff, workflow primitives for chains and fan-out, shared state, propagated cancellation and budgets, and human-in-the-loop pause and resume.

Phase 5, observability and release. tracing instrumentation, cost accounting, record and replay for deterministic tests, and a mock provider for testing agents without network access.

Out of scope: prompt-template DSLs, a built-in vector database, a web UI or server, and fine-tuning orchestration. Freyja is a library, not a platform.

Also out of scope: automatic retries. Backing off means sleeping, and Freyja exposes async fn without spawning so the caller picks the runtime; retrying internally would take that choice away to save ten lines. Error::is_retryable() and retry_after() make the decision cheap instead, and compose with backon or tower::retry. See Errors.

Development

cargo test
cargo clippy --all-targets -- -D warnings
cargo fmt --check

Rust edition 2024, minimum toolchain 1.88, verified in CI. tokio and dotenvy are dev-dependencies used only by the examples, so a consumer does not inherit them.

Contributions: Architecture explains the layout, Capability model explains what Freyja is allowed to refuse and why that is almost nothing, Adding a dialect covers new wire formats, and reaching a new vendor usually needs no code at all, see Custom providers.

License

MIT. See LICENSE.

About

A provider-neutral LLM client for Rust: one request model, four wire dialects, any compatible endpoint.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages