ai4RAG is an optimization engine for RAG Templates that is LLM and vector database provider-agnostic.
It accepts a variety of RAG Templates and a search space definition, then returns an initialized RAG Template with optimal parameter values (called a RAG Pattern).
Important
ai4rag is provider-agnostic. It reaches foundation and embedding models through the stock openai SDK, so any OpenAI-compatible endpoint works — a hosted API, a self-managed server (vLLM, TGI, Ollama, …), or an OpenShift AI Models-as-a-Service (MaaS) deployment, the integration ai4rag ships helpers for out of the box. You can also plug in your own foundation model, embedding model, or vector store by implementing the matching Base* interface.
To run an experiment you'll need one foundation model and one embedding model (from any of the above), plus a vector store (remote Milvus, embedded Milvus Lite, or PostgreSQL/pgvector) connected directly via ai4rag.rag.vector_store.
ai4rag reaches foundation and embedding models through the stock openai SDK, so it works with any OpenAI-compatible endpoint — a hosted API, a self-managed server (vLLM, TGI, Ollama, …), or an OpenShift AI Models-as-a-Service (MaaS) deployment. Prefer something else entirely? Implement BaseFoundationModel / BaseEmbeddingModel and pass your own models straight into an experiment.
MaaS is the integration ai4rag ships helpers for, so the walkthrough below uses it:
- SDK: openai >= 2, < 3 (Python package used by ai4RAG; installs with this project).
- Deployment: an OpenShift AI MaaS instance exposing at least one foundation model and one embedding model.
- Endpoints: MaaS serves everything from a single OpenAI-compatible endpoint —
MAAS_BASE_URL(host-only or/v1-suffixed; normalized automatically). One client lists the available models (models.list()) and serves chat/completions and embeddings for all of them. Model ids are used verbatim, exactly asmodels.list()reports them.
Features used by ai4rag
When using the MaaS backend, ai4rag relies on:
- Embeddings — Text embeddings via the
embeddingsendpoint (e.g. for indexing and query encoding). Becausemodels.list()carries no metadata, embedding dimension and context length are auto-detected at construction (or supplied viaparams). - Chat / completions — Foundation model integration for answer generation when evaluating RAG patterns.
Vector storage is independent of MaaS: ai4rag connects directly to remote Milvus, embedded Milvus Lite, or PostgreSQL/pgvector via the config classes in ai4rag.rag.vector_store (see Vector stores below).
ai4RAG talks to the vector store directly through provider-specific clients — no MaaS deployment is required for this part. Pick a provider and pass its config to AI4RAGExperiment as vector_store_config:
MilvusConfig— remote Milvus server or Zilliz Cloud only.urimust be ahttp(s)://URL (TLS and self-signed CAs viaserver_cert); anything else (a bare host, a file path, an empty string) raisesValueError. This is a deliberate safety check: a mistyped or unreachableMILVUS_URInow fails loudly instead of silently falling back to a throwaway local database. Supports hybrid search (dense + BM25).MilvusLiteConfig— the embedded, zero-server Milvus Lite engine, backed by a localdb_pathfile (default"./ai4rag_milvus_lite.db") — no setup required, ideal for local development and small-scale workloads. Also supports hybrid search (dense + BM25); rejectshttp(s)://values (useMilvusConfigfor those).PGVectorConfig— PostgreSQL with thepgvectorextension. Hybrid search (dense +tsvectorfull-text).
Each config is a frozen dataclass with a .from_env() constructor and an env_vars attribute listing the environment variables it reads (e.g. MILVUS_URI for MilvusConfig, MILVUS_LITE_DB_PATH for MilvusLiteConfig, PGVECTOR_HOST for PGVectorConfig).
Note
Milvus Lite is intended for local development, tests, and small-scale workloads (prototyping, up to roughly 1M vectors) — not production serving. For production or large corpora, use a remote Milvus server (MilvusConfig), Zilliz Cloud, or pgvector.
ai4RAG uses docling-core for document representation and chunking. Documents are represented as DoclingDocument instances, and the DoclingChunker leverages docling's HybridChunker for structure-aware, token-aware chunking. docling-core, openai, and the vector store clients (pymilvus with Milvus Lite, pgvector, asyncpg) are all installed automatically with ai4rag.
If you are running ai4rag on a disconnected cluster (no internet access), you must pre-download the required ML models before execution.
ai4rag depends on models from two sources:
| Component | Size | Purpose | Environment Variable |
|---|---|---|---|
| Docling artifacts | ~300-400 MB | Document text extraction and optional OCR | DOCLING_ARTIFACTS_PATH |
| HuggingFace models | Variable | Embeddings and foundation models | HF_HOME |
-
On an internet-connected machine, download the artifacts:
pip install 'ai4rag[text-extraction]' # Trigger Docling model download python -c "from docling.document_converter import DocumentConverter; \ converter = DocumentConverter(); \ converter.convert_document_string('/tmp/test.txt')" # Pre-download HuggingFace models export HF_HOME=/path/to/hf_cache python -c "from transformers import AutoTokenizer; \ AutoTokenizer.from_pretrained('BAAI/bge-m3')"
-
Transfer the artifacts to your disconnected cluster:
rsync -av ~/.cache/docling/ cluster:/offline/docling/ rsync -av /path/to/hf_cache/ cluster:/offline/hf_cache/ -
On the disconnected cluster, set environment variables:
export DOCLING_ARTIFACTS_PATH=/offline/docling export HF_HOME=/offline/hf_cache export HF_HUB_OFFLINE=1 # Enforce offline mode
The provided maas_indexing_template.ipynb includes a "Prerequisites for Disconnected Clusters" section with:
- Validation cells to verify artifact availability
- Configuration examples for custom model paths
- Step-by-step download instructions in the appendix
All indexing notebooks automatically detect and use pre-downloaded artifacts when DOCLING_ARTIFACTS_PATH is set.
- Prepare a MaaS client to integrate with your models.
- Prepare your knowledge base documents for the experiment.
- Prepare
benchmark_data.jsonwith evaluation questions and answers. - Define and constrain your search space.
- Configure the optimizer.
- Create and run the experiment.
To enable full integration with MaaS, build a single client that lists the available models and serves them all — ai4rag reuses it for every foundation and embedding model wrapper.
The dev_utils helper create_dev_maas_client() reads MAAS_BASE_URL / MAAS_API_KEY and builds that client for you.
Tip
Store your credentials securely in a .env file.
from dotenv import load_dotenv, find_dotenv
from dev_utils.utils import create_dev_maas_client
load_dotenv(find_dotenv())
client = create_dev_maas_client() # reads MAAS_BASE_URL / MAAS_API_KEYNote
dev_utils is only available when cloning the repository. For the equivalent setup using the
public API (the single OpenAI client built with create_maas_client),
see the Provider-Agnostic Design guide.
Prepare a set of documents to serve as the knowledge base for retrieval.
Documents are represented as DoclingDocument instances (from the docling-core library).
A local folder is all you need — ai4rag does not require object storage.
Convert the folder with Docling, naming each document by its path relative to that folder:
from pathlib import Path
from docling.document_converter import DocumentConverter
documents_root = Path("<path to the documents folder>")
converter = DocumentConverter()
documents = []
for file_path in sorted(p for p in documents_root.rglob("*") if p.is_file()):
document = converter.convert(file_path).document
# The document's key: what benchmark data references.
document.name = str(file_path.relative_to(documents_root))
documents.append(document)Important
Each document's name is its key — the identifier carried through chunking, indexing and evaluation,
and the value correct_answer_document_keys must reference. Set it explicitly: Docling otherwise derives a
name from the file stem, so two files named setup.pdf in different folders would collide.
Already keeping your corpus in a bucket? discover_documents() and extract_text() name each document by
its full S3 object key, so that is what correct_answer_document_keys must reference — prefix included.
Create a benchmark_data.json file following this schema:
[
{
"question": "<question_1>",
"correct_answers": [
"<answer 1 for question 1>",
"<answer 2 for question 1>"
],
"correct_answer_document_keys": ["<list of document keys based on which correct answers were generated>"]
},
{
"question": "<question_2>",
"correct_answers": [
"<answer 1 for question 2>",
"<answer 2 for question 2>"
],
"correct_answer_document_keys": ["<list of document keys based on which correct answers were generated>"]
}
]All benchmark questions and answers must be derived from your knowledge base documents, and every
correct_answer_document_keys entry must equal the name of one of the documents you loaded above.
import pandas as pd
benchmark_data = pd.read_json("<path to benchmark_data.json>")The search space defines all possible parameter combinations, where each combination creates a unique RAG Pattern. During the experiment, the engine will optimize the RAG Pattern for the selected metric over the given search space, using an objective function to evaluate each configuration.
from ai4rag.search_space.src.parameter import Parameter
from ai4rag.search_space.src.search_space import AI4RAGSearchSpace
from dev_utils.utils import build_maas_model
search_space = AI4RAGSearchSpace(
params=[
Parameter(
name="foundation_model",
param_type="C",
values=[build_maas_model(client, model_id="qwen3-8b-fp8-dynamic", model_type="llm")],
),
Parameter(
name="embedding_model",
param_type="C",
values=[
build_maas_model(
client,
model_id="bge-m3",
model_type="embedding",
embedding_params={"embedding_dimension": 1024, "context_length": 8192},
)
],
),
Parameter(
name="chunking_method",
param_type="C",
values=["recursive", "hybrid"],
),
Parameter(
name="chunk_size",
param_type="C",
values=[512, 1024, 2048],
),
Parameter(
name="chunk_overlap",
param_type="C",
values=[0, 128, 256],
),
]
)Tip
chunking_method controls the chunking strategy: "recursive" uses LangChain's RecursiveCharacterTextSplitter, while "hybrid" uses docling's structure-aware HybridChunker (requires chunk_overlap=0).
When omitted, both methods are included by default.
Tip
To validate model IDs and build a search space from a MaaS deployment in one call, use prepare_search_space_with_maas() from ai4rag.search_space.prepare, passing the MaaS client and the foundation/embedding model IDs per type.
You have full control over the optimization algorithm. Configure the GAMOptimizer by adjusting GAMOptSettings.
from ai4rag.core.hpo.gam_opt import GAMOptSettings
optimizer_settings = GAMOptSettings(
max_evals=10, n_random_nodes=4
)Using the information from the previous steps, create an experiment and run the ai4rag optimization engine.
Note
Select the vector store by passing a vector_store_config to AI4RAGExperiment:
MilvusLiteConfig() (or MilvusLiteConfig(db_path="./ai4rag.db")) for a zero-config, local Milvus Lite store, or
MilvusConfig.from_env() / PGVectorConfig.from_env() for a server-backed store. All support hybrid (dense + keyword) search.
from ai4rag.core.experiment.experiment import AI4RAGExperiment
from ai4rag.rag.vector_store import MilvusConfig
from ai4rag.utils.event_handler import LocalEventHandler
experiment = AI4RAGExperiment(
documents=documents,
benchmark_data=benchmark_data,
search_space=search_space,
vector_store_config=MilvusConfig.from_env(),
optimizer_settings=optimizer_settings,
event_handler=LocalEventHandler(output_path="<local-path-to-store-your-output-files>"),
)
experiment.search()
best_eval = experiment.results.get_best_evaluations(k=1)[0]
print(best_eval)
print(f"Best pattern: {best_eval.pattern_name} (score: {best_eval.final_score})")Note
Each trial closes its vector store once it finishes, so EvaluationResult no longer exposes a reusable rag_pattern. Read the outcome from its fields (pattern_name, final_score, scores, rag_params); rebuild the pattern from those settings if you want to run inference.
Tip
For production use, implement your own custom EventHandler to handle status changes and artifacts produced during the experiment.
See the BaseEventHandler implementation for reference.
Pull requests are very welcome! Make sure your patches are well tested. Ideally create a topic branch for every separate change you make.
This project uses uv for dependency management.
# Clone the repository
git clone https://github.com/IBM/ai4rag.git
cd ai4rag
# Install all development dependencies
uv sync --extra dev
# Run tests
uv run pytest tests/unit/
# Check code style
uv run black --check ai4rag/
uv run pylint ai4rag/
# Build and serve documentation locally
uv run mkdocs serve- Fork the repo
- Create your feature branch (
git checkout -b my-new-feature) - Commit your changes (
git commit -s -am 'Added some feature') - Push to the branch (
git push origin my-new-feature) - Create new Pull Request
See more details in contributing section.