Skip to content

Evaluation

Evaluation lets you systematically measure how well your BeeAI agents perform on tasks like multi-hop question answering, tool usage, and factual accuracy. Rather than manually inspecting outputs, you define a dataset of questions with expected answers and run automated metrics to score agent responses.

BeeAI integrates with two popular evaluation frameworks via dedicated adapters:

  • DeepEval: evaluation framework with built-in and custom metrics
  • Ragas: async evaluation framework focused on RAG and agent pipelines

The adapters bridge BeeAI’s ChatModel to each evaluation framework’s LLM interface, so metrics can use BeeAI-powered models for judging agent outputs without requiring separate API configurations.

Wraps a BeeAI ChatModel as a DeepEvalBaseLLM, enabling DeepEval metrics to call any BeeAI-supported model for LLM-as-judge evaluations. Supports structured output via Pydantic schemas.

from beeai_framework.evaluation.adapters import DeepEvalLLM
# Create from a model name string
eval_llm = DeepEvalLLM.from_name("openai:gpt-4o")
# Or wrap an existing ChatModel
from beeai_framework.backend import ChatModel
model = ChatModel.from_name("ollama:llama3.1:8b")
eval_llm = DeepEvalLLM(model)

Bridges Ragas directly to BeeAI’s ChatModel. Handles both async and sync structured output generation.

from beeai_framework.evaluation.adapters import InstructorRagasLLM
# Create from a model name string
eval_llm = InstructorRagasLLM.from_name("openai:gpt-4o")
# Or wrap an existing ChatModel
from beeai_framework.backend import ChatModel
model = ChatModel.from_name("ollama:llama3.1:8b")
eval_llm = InstructorRagasLLM(model)

The rest of this page uses one reference example throughout: a multi-hop QA agent evaluated against a small bundled dataset. Evaluation experiments require two components: a dataset of test cases and an agent to evaluate.

The bundled sample dataset is built from real HotpotQA records: multi-hop questions that can’t be answered from a single source. The agent must retrieve and combine facts from two or more separate Wikipedia articles (e.g. comparing which of two magazines launched first, or tracing a company name to its headquarters city).

HotpotQA’s own records carry a context field (every sentence from each candidate article, including irrelevant “distractor” paragraphs) and a supporting_facts field (pointers, as title + sent_id, into context, marking which sentences actually justify the answer). This example’s dataset simplifies that structure for direct use by the agent and metrics:

HotpotQA fieldOur fieldHow it’s derived
question, answerquestion, expected_answerCopied as-is
supporting_facts (title + sent_id pointers)supporting_sentencesEach pointer is resolved against context and the actual sentence text is copied in, instead of kept as an index
supporting_facts.title (deduplicated)supporting_titlesThe distinct article titles the supporting facts point to
(no HotpotQA equivalent)expected_tool_callsCount of distinct supporting_titles; how many Wikipedia lookups the agent is expected to make
context’s distractor paragraphs, type, level, id(dropped)Not needed for this example’s simpler dataset

The dataset is a JSON file where each item contains:

FieldDescription
questionThe input question for the agent
expected_answerThe ground-truth answer
supporting_sentencesRelevant facts the agent should reference
supporting_titlesSource titles for the supporting context
expected_tool_callsExpected number of Wikipedia tool calls
# dataset.py: load evaluation items
import json
from pathlib import Path
def load_items() -> list[dict]:
with open(Path(__file__).parent / "dataset.json", encoding="utf-8") as f:
return json.load(f)

The evaluation agent is a RequirementAgent configured for multi-hop QA with structured JSON output. It uses WikipediaTool and OpenMeteoTool to gather information, plus a minimal CalculatorTool for arithmetic. CalculatorTool is a dependency-free tool that evaluates expressions through a restricted AST walk, so it can’t execute arbitrary code and needs no external interpreter server.

from beeai_framework.agents.requirement import RequirementAgent
from beeai_framework.backend import ChatModel
from beeai_framework.memory import UnconstrainedMemory
from beeai_framework.tools.search.wikipedia import WikipediaTool
from beeai_framework.tools.weather import OpenMeteoTool
from examples.evaluation.calculator_tool import CalculatorTool
agent = RequirementAgent(
llm=ChatModel.from_name("ollama:llama3.1:8b"),
tools=[WikipediaTool(), OpenMeteoTool(), CalculatorTool()],
memory=UnconstrainedMemory(),
role="You are an expert Multi-hop QA agent.",
instructions=["Answer in JSON format with: answer, tool_used, supporting_titles, supporting_sentences, reasoning_explanation"],
)
VariablePurposeDefault
AGENT_CHAT_MODEL_NAMEModel used by the agent under testollama:llama3.1:8b
EVAL_CHAT_MODEL_NAMEModel used by LLM-as-judge metricsFalls back to agent model
EVAL_LOG_LLM_CALLSSet to "true" to log evaluation LLM callsDisabled

DeepEval experiments run as async Python scripts. Each dataset item becomes a test case scored against multiple metrics:

  • Answer quality: AnswerRelevancyMetric, AnswerLLMJudgeMetric, ExactMatchMetric, FaithfulnessMetric, ContextualRecallMetric
  • Tool usage: ToolCorrectnessMetric, ToolUsageMetric, ArgumentCorrectnessMetric
  • Facts: FactsSimilarityMetric
Terminal window
cd python/examples/evaluation
EVAL_CHAT_MODEL_NAME=openai:gpt-4o python deepeval/experiment.py

A pickle of the raw DeepEval results is written next to experiment.py as eval_results_raw.pkl. The pass/fail summary is printed to the log.

Ragas experiments run as async Python scripts with metrics computed per dataset item:

  • Context: ContextPrecision, ContextRecall
  • Answer: ExactMatch, AnswerAccuracy
  • Tool usage: ToolCallAccuracy
  • Facts: FactsSimilarityMetric
Terminal window
cd python/examples/evaluation/ragas
python dataset.py # build the dataset (once; no module-load side effects)
EVAL_CHAT_MODEL_NAME=openai:gpt-4o python experiment.py

dataset.py exposes its work via build_dataset(); importing the module is a no-op, so it’s safe to import the helper from other scripts.

Each Ragas run writes two files into ragas/data/experiments/, both overwritten on every run:

FilePurpose
my_results.csvTabular results (written by Ragas via arun(name="my_results"))
my_results.pklPickle of the raw Ragas Experiment object (for reloading)

The stable CSV filename comes from passing name="my_results" to my_experiment.arun(...). Without it, Ragas auto-generates a new random filename (e.g. dazzling_minsky.csv) per run.

Both frameworks support custom metrics that extend their base classes. The evaluation examples include three custom metrics:

MetricWhat it measuresScore
AnswerLLMJudgeMetricSemantic similarity between actual and expected answers, judged by an LLM0–1
ToolUsageMetricWhether the agent called the expected tools the expected number of times0–1
FactsSimilarityMetricSemantic similarity of supporting facts between actual and expected outputs0–1

To create your own custom metric, extend the appropriate base class (DeepEvalBaseLLM or Ragas BaseMetric) and use the adapter to call BeeAI models for LLM-based judging.