Description
TruLens is an open source tool designed to evaluate and trace AI agents, providing insights into their performance and efficiency. By leveraging OpenTelemetry, TruLens instruments AI agents to capture detailed traces of their operations. Each step is scored using benchmarked LLM judges, offering a clear understanding of where an agent may fail and where costs can be optimized without losing quality.
The tool is particularly useful for developers looking to move from subjective assessments to objective metrics. TruLens records latency, inputs, outputs, tokens, and costs per step, allowing developers to trace the cause of any suboptimal performance. The judges not only provide scores but also explain them, highlighting specific areas for improvement.
TruLens supports a wide range of evaluations, including tool selection, plan adherence, execution efficiency, and more. It also assesses the groundedness and relevance of answers, ensuring that AI agents provide accurate and contextually appropriate responses. For developers working with MCP apps, TruLens evaluates tool calling, quality, and span tracing.
The tool is flexible, allowing users to start with state-of-the-art judges and then align them to their specific domain using private data. This customization is facilitated by adding rubrics, examples, and adjusting score ranges. TruLens also supports quickstarts for popular frameworks like LangChain, LangGraph, and LlamaIndex, making it accessible for a wide range of applications.
TruLens was originally created by TruEra and is now maintained by Snowflake, ensuring ongoing support and development. Its open-source nature means there is no lock-in, and developers can integrate it with their existing stack without needing to rewrite their applications.
TruLens: Evals and Tracing for AI Agents's Core Features
OpenTelemetry instrumentation
Benchmarked LLM judges
Trace latency, inputs, outputs
Score explanation
Customizable judges
Tool selection evaluation
Plan adherence assessment
Execution efficiency analysis
Groundedness and relevance checks
MCP span tracing
How to use TruLens: Evals and Tracing for AI Agents?
Install: Use pip to install TruLens
Instrument: Add OpenTelemetry to your AI agent
Evaluate: Run evaluations with LLM judges
Optimize: Adjust based on trace and score feedback
TruLens: Evals and Tracing for AI Agents's Use Cases
- AI Agent Evaluation
- Cost Optimization
- Custom Judge Alignment
- Tool Selection Analysis
- Plan Adherence Assessment







