docs/docs/genai/eval-monitor/index.mdx
import Tabs from "@theme/Tabs" import TabItem from "@theme/TabItem" import ImageBox from "@site/src/components/ImageBox"; import TabsWrapper from "@site/src/components/TabsWrapper"; import TilesGrid from "@site/src/components/TilesGrid"; import TileCard from "@site/src/components/TileCard"; import { Rocket, Scale } from "lucide-react"; import useBaseUrl from '@docusaurus/useBaseUrl'; import GenAIDemoCard from "@site/src/content/genai_demo_card.mdx";
MLflow's evaluation and monitoring capabilities help you systematically measure, improve, and maintain the quality of your LLM applications and AI agents throughout their lifecycle from development through production.
<video src={useBaseUrl("/images/mlflow-3/eval-monitor/evaluation-result-video.mp4")} controls loop autoPlay muted aria-label="Prompt Evaluation" />
<GenAIDemoCard />A core tenet of MLflow's evaluation capabilities is Evaluation-Driven Development. This is an emerging practice to tackle the challenge of building high-quality LLM/Agentic applications. MLflow is an open source AI engineering platform that is designed to support this practice and help you quickly build production-quality AI agents and LLM applications.
<ImageBox src="/images/mlflow-3/eval-monitor/evaluation-driven-development.png" alt="Evaluation Driven Development" width="90%"/> #### Create and maintain a High-Quality Dataset
Before you can evaluate your LLM application or AI agent, you need test data. **Evaluation Datasets** provide a centralized repository for managing test cases, ground truth expectations, and evaluation data at scale.
Think of Evaluation Datasets as your "test database" - a single source of truth for all the data needed to evaluate your AI systems. They transform ad-hoc testing into systematic quality assurance.
[Learn more →](/genai/datasets)
</div>
<div class="flex-item padding-md">

</div>
</div>
</div>
#### Track Annotation and Human Feedbacks
Human feedback is essential for building high-quality LLM applications and AI agents that meet user expectations. MLflow supports collecting, managing, and utilizing feedback from end-users and domain experts.
Feedbacks are attached to traces and recorded with metadata, including user, timestamp, revisions, etc.
[Learn more →](/genai/assessments/feedback)
</div>
<div class="flex-item padding-md">

</div>
</div>
</div>
#### Scale Quality Assessment with Automation
Quality assessment is a critical part of building high-quality LLM applications and AI agents, however, it is often time-consuming and requires human expertise. LLMs are powerful tools to automate quality assessment.
MLflow offers various built-in [LLM-as-a-Judge](https://mlflow.org/llm-evaluation) scorers to help automate the process, as well as a flexible toolset to build your own LLM judges with ease.
[Learn more →](/genai/eval-monitor)
</div>
<div class="flex-item padding-md">

</div>
</div>
</div>
#### Evaluate and Enhance quality
Systematically assessing and improving the quality of LLM applications and AI agents is a challenge. MLflow provides a comprehensive set of tools to help you evaluate and enhance the quality of your applications.
Being the industry's most-trusted open source [AI engineering platform](https://mlflow.org/genai) for agents and LLM applications, MLflow provides a strong foundation for tracking your evaluation results and effectively collaborating with your team.
[Learn more →](/genai/eval-monitor/quickstart)
</div>
<div class="flex-item padding-md">

</div>
</div>
</div>
#### Monitor Applications in Production
Understanding and optimizing LLM application and AI agent performance is crucial for efficient operations. [MLflow Tracing](https://mlflow.org/llm-tracing) captures key metrics like latency and token usage at each step, as well as various quality metrics, helping you identify bottlenecks, monitor efficiency, and find optimization opportunities.
[Learn more →](/genai/tracing/prod-tracing)
</div>
<div class="flex-item padding-md">

</div>
</div>
</div>
Each evaluation is defined by three components:
<table style={{ width: "100%" }}> <thead> <tr> <th style={{ width: "30%" }}>Component</th> <th style={{ width: "70%" }}>Example</th> </tr> </thead> <tbody> <tr> <td><strong>Dataset</strong> <small>Inputs & expectations (and optionally pre-generated outputs and traces)</small></td> <td> <pre><code>\[ \{"inputs": \{"question": "2+2"\}, "expectations": \{"answer": "4"\}}, \{"inputs": \{"question": "2+3"\}, "expectations": \{"answer": "5"\}\} \]</code></pre> </td> </tr> <tr> <td><strong>Scorer</strong> <small>Evaluation criteria</small></td> <td> <pre><code>@scorer def exact_match(expectations, outputs): return expectations == outputs</code></pre> </td> </tr> <tr> <td><strong>Predict Function</strong> <small>Generates outputs for the dataset</small></td> <td> <pre><code>def predict_fn(question: str) -> str: response = client.chat.completions.create( model="gpt-4o-mini", messages=\[\{"role": "user", "content": question\}\] ) return response.choices[0].message.content</code></pre> </td> </tr> </tbody> </table>The following example shows a simple evaluation of a dataset of questions and expected answers.
import os
import openai
import mlflow
from mlflow.genai.scorers import Correctness, Guidelines
client = openai.OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
# 1. Define a simple QA dataset
dataset = [
{
"inputs": {"question": "Can MLflow manage prompts?"},
"expectations": {"expected_response": "Yes!"},
},
{
"inputs": {"question": "Can MLflow create a taco for my lunch?"},
"expectations": {"expected_response": "No, unfortunately, MLflow is not a taco maker."},
},
]
# 2. Define a prediction function to generate responses
def predict_fn(question: str) -> str:
response = client.chat.completions.create(
model="gpt-4o-mini", messages=[{"role": "user", "content": question}]
)
return response.choices[0].message.content
# 3.Run the evaluation
results = mlflow.genai.evaluate(
data=dataset,
predict_fn=predict_fn,
scorers=[
# Built-in LLM judge
Correctness(),
# Custom criteria using LLM judge
Guidelines(name="is_english", guidelines="The answer must be in English"),
],
)
Open the MLflow UI to review the evaluation results. You can use the following command to start the UI:
mlflow server --port 5000
You should see a new evaluation run is created under the "Runs" tab. Click on the run name to view the evaluation results.
<ImageBox src="/images/mlflow-3/eval-monitor/quickstart-eval-hero.png" alt="Evaluation Results" />