Evaluation
Overview
Section titled “Overview”Evaluation lets you systematically measure how well your BeeAI agents perform on tasks like multi-hop question answering, tool usage, and factual accuracy. Rather than manually inspecting outputs, you define a dataset of questions with expected answers and run automated metrics to score agent responses.
BeeAI integrates with two popular evaluation frameworks via dedicated adapters:
- DeepEval: evaluation framework with built-in and custom metrics
- Ragas: async evaluation framework focused on RAG and agent pipelines
Adapters
Section titled “Adapters”The adapters bridge BeeAI’s ChatModel to each evaluation framework’s LLM interface, so metrics can use BeeAI-powered models for judging agent outputs without requiring separate API configurations.
DeepEvalLLM
Section titled “DeepEvalLLM”Wraps a BeeAI ChatModel as a DeepEvalBaseLLM, enabling DeepEval metrics to call any BeeAI-supported model for LLM-as-judge evaluations. Supports structured output via Pydantic schemas.
from beeai_framework.evaluation.adapters import DeepEvalLLM
# Create from a model name stringeval_llm = DeepEvalLLM.from_name("openai:gpt-4o")
# Or wrap an existing ChatModelfrom beeai_framework.backend import ChatModelmodel = ChatModel.from_name("ollama:llama3.1:8b")eval_llm = DeepEvalLLM(model)InstructorRagasLLM
Section titled “InstructorRagasLLM”Bridges Ragas directly to BeeAI’s ChatModel. Handles both async and sync structured output generation.
from beeai_framework.evaluation.adapters import InstructorRagasLLM
# Create from a model name stringeval_llm = InstructorRagasLLM.from_name("openai:gpt-4o")
# Or wrap an existing ChatModelfrom beeai_framework.backend import ChatModelmodel = ChatModel.from_name("ollama:llama3.1:8b")eval_llm = InstructorRagasLLM(model)Example
Section titled “Example”The rest of this page uses one reference example throughout: a multi-hop QA agent evaluated against a small bundled dataset. Evaluation experiments require two components: a dataset of test cases and an agent to evaluate.
Dataset
Section titled “Dataset”The bundled sample dataset is built from real HotpotQA records: multi-hop questions that can’t be answered from a single source. The agent must retrieve and combine facts from two or more separate Wikipedia articles (e.g. comparing which of two magazines launched first, or tracing a company name to its headquarters city).
HotpotQA’s own records carry a context field (every sentence from each candidate article, including irrelevant “distractor” paragraphs) and a supporting_facts field (pointers, as title + sent_id, into context, marking which sentences actually justify the answer). This example’s dataset simplifies that structure for direct use by the agent and metrics:
| HotpotQA field | Our field | How it’s derived |
|---|---|---|
question, answer | question, expected_answer | Copied as-is |
supporting_facts (title + sent_id pointers) | supporting_sentences | Each pointer is resolved against context and the actual sentence text is copied in, instead of kept as an index |
supporting_facts.title (deduplicated) | supporting_titles | The distinct article titles the supporting facts point to |
| (no HotpotQA equivalent) | expected_tool_calls | Count of distinct supporting_titles; how many Wikipedia lookups the agent is expected to make |
context’s distractor paragraphs, type, level, id | (dropped) | Not needed for this example’s simpler dataset |
The dataset is a JSON file where each item contains:
| Field | Description |
|---|---|
question | The input question for the agent |
expected_answer | The ground-truth answer |
supporting_sentences | Relevant facts the agent should reference |
supporting_titles | Source titles for the supporting context |
expected_tool_calls | Expected number of Wikipedia tool calls |
# dataset.py: load evaluation itemsimport jsonfrom pathlib import Path
def load_items() -> list[dict]: with open(Path(__file__).parent / "dataset.json", encoding="utf-8") as f: return json.load(f)The evaluation agent is a RequirementAgent configured for multi-hop QA with structured JSON output. It uses WikipediaTool and OpenMeteoTool to gather information, plus a minimal CalculatorTool for arithmetic. CalculatorTool is a dependency-free tool that evaluates expressions through a restricted AST walk, so it can’t execute arbitrary code and needs no external interpreter server.
from beeai_framework.agents.requirement import RequirementAgentfrom beeai_framework.backend import ChatModelfrom beeai_framework.memory import UnconstrainedMemoryfrom beeai_framework.tools.search.wikipedia import WikipediaToolfrom beeai_framework.tools.weather import OpenMeteoToolfrom examples.evaluation.calculator_tool import CalculatorTool
agent = RequirementAgent( llm=ChatModel.from_name("ollama:llama3.1:8b"), tools=[WikipediaTool(), OpenMeteoTool(), CalculatorTool()], memory=UnconstrainedMemory(), role="You are an expert Multi-hop QA agent.", instructions=["Answer in JSON format with: answer, tool_used, supporting_titles, supporting_sentences, reasoning_explanation"],)Environment Variables
Section titled “Environment Variables”| Variable | Purpose | Default |
|---|---|---|
AGENT_CHAT_MODEL_NAME | Model used by the agent under test | ollama:llama3.1:8b |
EVAL_CHAT_MODEL_NAME | Model used by LLM-as-judge metrics | Falls back to agent model |
EVAL_LOG_LLM_CALLS | Set to "true" to log evaluation LLM calls | Disabled |
Running Evaluations
Section titled “Running Evaluations”DeepEval
Section titled “DeepEval”DeepEval experiments run as async Python scripts. Each dataset item becomes a test case scored against multiple metrics:
- Answer quality: AnswerRelevancyMetric, AnswerLLMJudgeMetric, ExactMatchMetric, FaithfulnessMetric, ContextualRecallMetric
- Tool usage: ToolCorrectnessMetric, ToolUsageMetric, ArgumentCorrectnessMetric
- Facts: FactsSimilarityMetric
cd python/examples/evaluationEVAL_CHAT_MODEL_NAME=openai:gpt-4o python deepeval/experiment.pyOutput files
Section titled “Output files”A pickle of the raw DeepEval results is written next to experiment.py as eval_results_raw.pkl. The pass/fail summary is printed to the log.
Ragas experiments run as async Python scripts with metrics computed per dataset item:
- Context: ContextPrecision, ContextRecall
- Answer: ExactMatch, AnswerAccuracy
- Tool usage: ToolCallAccuracy
- Facts: FactsSimilarityMetric
cd python/examples/evaluation/ragaspython dataset.py # build the dataset (once; no module-load side effects)EVAL_CHAT_MODEL_NAME=openai:gpt-4o python experiment.pydataset.py exposes its work via build_dataset(); importing the module is a no-op, so it’s safe to import the helper from other scripts.
Output files
Section titled “Output files”Each Ragas run writes two files into ragas/data/experiments/, both overwritten on every run:
| File | Purpose |
|---|---|
my_results.csv | Tabular results (written by Ragas via arun(name="my_results")) |
my_results.pkl | Pickle of the raw Ragas Experiment object (for reloading) |
The stable CSV filename comes from passing name="my_results" to my_experiment.arun(...). Without it, Ragas auto-generates a new random filename (e.g. dazzling_minsky.csv) per run.
Custom Metrics
Section titled “Custom Metrics”Both frameworks support custom metrics that extend their base classes. The evaluation examples include three custom metrics:
| Metric | What it measures | Score |
|---|---|---|
| AnswerLLMJudgeMetric | Semantic similarity between actual and expected answers, judged by an LLM | 0–1 |
| ToolUsageMetric | Whether the agent called the expected tools the expected number of times | 0–1 |
| FactsSimilarityMetric | Semantic similarity of supporting facts between actual and expected outputs | 0–1 |
To create your own custom metric, extend the appropriate base class (DeepEvalBaseLLM or Ragas BaseMetric) and use the adapter to call BeeAI models for LLM-based judging.