About this postCo-edited with Claude (Anthropic).
Summary
- Evaluating GenAI apps splits into two tracks: quality evaluation, which checks accuracy and RAG faithfulness, and red teaming, which hunts for jailbreaks and data leaks. DeepEval covers the former, DeepTeam the latter
- DeepEval is an evaluation library that runs on pytest, with 50+ metrics. It covers RAG, agents, conversations, MCP, and even voice — the broadest range in open source. Its weak points are the cost and score variance of LLM-as-judge, plus frequent breaking changes (4.2.0 flipped the direction of some scores)
- DeepTeam is a red-teaming tool that hooks into your app through a single Python callback, with presets for OWASP LLM Top 10, the EU AI Act, and more. But its docs and code disagree in noticeable ways — even the default models differ
- Alternatives: for quality evaluation, promptfoo, Langfuse, Opik, Ragas, Inspect AI and others; for red teaming, promptfoo, PyRIT, garak, Giskard and others. DeepEval / DeepTeam have no production monitoring, so pairing them with Langfuse or Opik for production is the realistic setup
- What surprised me most while looking for alternatives was how hard the industry is being reshuffled. promptfoo was acquired by OpenAI, Langfuse by ClickHouse, Arize is being acquired by Dynatrace, Galileo went to Cisco, and OpenAI Evals shuts down on 2026-11-30. When choosing a tool, check not just the features but “who owns it now, and will it stay vendor-neutral?”
- As an idea outside the research: judge costs might be cut with a local LLM. A company that already owns hardware like a DGX Spark could use an open frontier model as the judge, saving API spend while also pinning the judge and keeping data in-house. Calibration is still mandatory
Introduction
I read 『生成AIアプリケーション評価入門』 (roughly “Introduction to Evaluating Generative AI Applications”, Shinsuke Matsuki, Gijutsu-Hyoronsha) — a Japanese book. It starts from classic evaluation metrics like the confusion matrix and BLEU / ROUGE, moves through the ISO/IEC 25059 quality model and the OWASP LLM Top 10, and finishes with a hands-on part that uses DeepEval and DeepTeam for evaluation and red teaming. As a systematic introduction it was easy to read and not too thick, which made it just right for getting the big picture of evaluation.
* Contains Amazon Associates links
After finishing the book, two things bugged me. First, how far can DeepEval and DeepTeam actually go, and where are their weak spots? Second, are there other options? DeepEval can’t be the only evaluation tool out there, and without comparing, I can’t tell where DeepEval stands.
So I dug into DeepEval and DeepTeam through their official docs and the source code of the published packages, and then went exploring for alternatives. As I searched, something caught my eye before any feature differences did: this industry was being reshuffled at a furious pace in 2026. Roughly half of the major tools were acquired, shut down their service, or stalled in development within the past year.
Note that I haven’t run DeepEval / DeepTeam myself. This is a summary of what I learned from the book plus the official docs, PyPI, GitHub, and source code. Versions, star counts, and prices are all as of 2026-10-05.
Evaluation splits into two tracks
To set things up first: “evaluation” of GenAI apps broadly splits into two. The book also covers DeepEval and DeepTeam in separate chapters.
| Track | What it checks | Covered by | Main competitors |
|---|---|---|---|
| Quality evaluation (Eval) | Accuracy, faithfulness, RAG retrieval quality, agent task completion, etc. | DeepEval | Ragas, promptfoo, Langfuse, LangSmith, Opik, MLflow, Braintrust, Inspect AI, cloud evaluation services |
| Safety evaluation (Red Teaming) | Jailbreaks, prompt injection, data leakage, excessive agency, etc. | DeepTeam | promptfoo, PyRIT, garak, Giskard, commercial (Prisma AIRS, Lakera) |
Both are open source from the same company, Confident AI, and DeepTeam depends on DeepEval at runtime. Below, I dig into both first, then look at alternatives for each.
Digging into DeepEval
Overview
| Item | Details |
|---|---|
| Developer | Confident AI |
| Positioning | ”Pytest for LLM apps”. Metrics run locally1 |
| License | Apache-2.02 |
| Latest | Python 4.2.8 (2026-10-02). Python 3.9+2. A TypeScript SDK also exists |
| GitHub | ~18.6k stars, releases almost weekly13 |
Version 4.0 (2026-05) added an evaluation harness for coding agents such as Claude Code / Cursor / Codex3. You can see from the tooling side, too, that the evaluation target is shifting from “chatbots” to “agents”.
Basic usage
There are only a few core concepts45.
- LLMTestCase: a single input/output exchange. Has
input,actual_output,expected_output,retrieval_context,tools_called, etc. - ConversationalTestCase: a multi-turn conversation.
ConversationSimulatorcan generate conversations automatically - Golden / EvaluationDataset: test data that holds only the expectations. The app is run at evaluation time to fill in the outputs
evaluate(): the entry point for evaluating from a scriptassert_test()anddeepeval test run: the entry point for running on pytest in CI. Has flags for parallel runs (-n), caching (-c), and repeats (-r)6@observe(): captures traces and applies metrics per span or to the whole trace
Following the official style, a test looks like this (as of 4.2.x).
# test_app.py → `deepeval test run test_app.py`from deepeval import assert_testfrom deepeval.metrics import GEval, AnswerRelevancyMetric, FaithfulnessMetricfrom deepeval.test_case import LLMTestCase, SingleTurnParams
correctness = GEval( name="Correctness", criteria="Determine if the 'actual output' is correct based on the 'expected output'.", evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT], threshold=0.5,)
def test_refund(): tc = LLMTestCase( input="What if these shoes don't fit?", actual_output="You have 30 days to get a full refund at no extra cost.", expected_output="We offer a 30-day full refund at no extra costs.", retrieval_context=["All customers are eligible for a 30 day full refund at no extra costs."], ) assert_test(tc, [correctness, AnswerRelevancyMetric(threshold=0.7), FaithfulnessMetric()])To evaluate an agent, combine it with tracing.
from deepeval.tracing import observefrom deepeval.dataset import EvaluationDataset, Goldenfrom deepeval.metrics import TaskCompletionMetric
@observe()def agent(q): ...
ds = EvaluationDataset(goldens=[Golden(input="Plan a two-day trip to Paris")])for g in ds.evals_iterator(metrics=[TaskCompletionMetric()]): agent(g.input)DeepEval’s defining trait is that evaluations look just like ordinary unit tests, which lowers the psychological barrier to putting them in CI.
Metrics
The official docs advertise “50+“7. Here are the main ones by category.
| Category | Metrics |
|---|---|
| General / custom | G-Eval (criteria written in natural language), DAG (decision tree that pins down the judging steps), Arena G-Eval (pairwise comparison), custom metrics |
| RAG | Answer Relevancy, Faithfulness, Contextual Precision / Recall / Relevancy |
| Agents | Task Completion, Step Efficiency, Plan Adherence, Plan Quality (whole trajectory) / Tool Correctness, Argument Correctness (per component) |
| Conversation | Knowledge Retention, Role Adherence, Conversation Completeness, Turn Relevancy, etc. |
| Safety | Bias, Toxicity, PII Leakage, Misuse, Non-Advice, Role Violation |
| MCP | MCP Task Completion, MCP Use, Multi-Turn MCP Use |
| Multimodal | Image generation / editing evaluation, metrics for voice agents (4.1.10+) |
| Non-LLM | JSON Correctness, Exact Match, Pattern Match |
| Classifiers (4.2+) | Judges that return labels instead of scores. Refusal, Escalation, Prompt Injection, Data Leakage, etc.8 |
Things to know about how it works:
- Almost everything is LLM-as-judge. Each metric uses one of QAG (extract claims and judge them one by one), G-Eval, or DAG
- Scores range from 0 to 1 and come with a reason. The default threshold is 0.5;
strict_mode=Truemakes it a 0/1 judgment - In 4.2.0 every metric was unified to “higher is better”. Before that, Bias and Toxicity pointed the other way, so thresholds in old code need revisiting
- The default judge is OpenAI (
gpt-5.4in 4.2.8). Anthropic, Bedrock, Gemini, local OpenAI-compatible servers like Ollama and vLLM, LiteLLM, and more are built in, and you switch withdeepeval set-<provider>1
There’s also synthetic data generation that builds evaluation data (Goldens) from documents, benchmarks like MMLU and GSM8K, prompt optimization via GEPA / MIPROv2, and integrations with LangChain / LlamaIndex / OpenAI Agents SDK / Anthropic SDK and others910. On synthetic data, I found it honest that the official docs themselves recommend the order “human-curated data first, then production traffic, synthetic data last”9.
Data handling and the paid cloud
DeepEval itself is free, but team dashboards, online evaluation, and annotation live on the paid Confident AI side. Prices as of 2026-10-0511:
| Plan | Monthly | Main contents |
|---|---|---|
| Free | $0 | 2 seats, 1 project, 5 test runs per week |
| Starter | $200 | Unlimited seats, 5 projects, online evaluation, annotation queues |
| Team | $2,000 | Unlimited projects, versioning, RBAC, SSO |
| Enterprise | Custom quote | On-prem, data residency options (including Japan). The red-teaming module is “Enterprise++” |
It used to start at $19.99 per user per month, so the pricing structure has changed a lot.
Three things to watch in data handling12:
- Evaluation data is not sent to the cloud unless you run
deepeval loginor setCONFIDENT_API_KEY - However, anonymous telemetry is on by default (sent to PostHog). Stop it with
DEEPEVAL_TELEMETRY_OPT_OUT=1 - It auto-loads
.env/.env.localon import. In CI,DEEPEVAL_DISABLE_DOTENV=1is recommended
Strengths and weaknesses
Strengths:
- The broadest metric set in open source, all the way to agents, MCP, and voice
- Evaluations are written as code and go straight into CI via pytest
- Many judge model options, including fully local setups with Ollama or vLLM
- AWS’s Bedrock AgentCore Evaluations adopted DeepEval metrics as third-party evaluators13
Weaknesses:
- Cost and latency. QAG-style metrics call the judge many times per case, so as test counts grow, token costs and rate limits start to bite14
- Scores fluctuate. Scores shift when the judge model is updated, so you need to check against human labels before using them for release decisions14
- Lots of API changes. The rename from
LLMTestCaseParamstoSingleTurnParams, the score-direction flip in 4.2.0, the DAG overhaul, and so on. Old tutorials often don’t run - Only engineers can write evaluations. Team sharing, history, and production monitoring are on the paid side15
As someone who learned from a book, it’s worth remembering that code from the book won’t necessarily run as-is on the current version.
Digging into DeepTeam
Overview
| Item | Details |
|---|---|
| Positioning | ”Penetration testing for LLMs”. Auto-generates attacks to surface app vulnerabilities16 |
| License | Apache-2.017 |
| Latest | 1.0.9 (PyPI, 2026-08-12). Python 3.9 to below 3.1417 |
| GitHub | ~3.0k stars. Created 2025-03; first stable 1.0.0 on 2025-11-121617 |
| Relation to DeepEval | A separate package, but depends on DeepEval at runtime and reuses its model layer and test case types18 |
GitHub Releases has only three entries, and the v1.0.9 tag carries the 1.0.0 release notes, so treat PyPI as the source of truth for versions.
Basic usage
from deepteam import red_teamfrom deepteam.vulnerabilities import Bias, PIILeakagefrom deepteam.attacks.single_turn import PromptInjection, ROT13from deepteam.attacks.multi_turn import LinearJailbreakingfrom deepteam.test_case import RTTurn
async def model_callback(input: str, turns: list[RTTurn] | None = None): reply = await my_app.generate(input) # the app under test return RTTurn(role="assistant", content=reply) # returning a plain str also works
ra = red_team( model_callback=model_callback, vulnerabilities=[Bias(types=["race", "gender"]), PIILeakage(types=["direct_disclosure"])], attacks=[PromptInjection(weight=2), ROT13(), LinearJailbreaking(num_turns=3)], simulator_model="gpt-4o-mini", evaluation_model="gpt-4o", attacks_per_vulnerability_type=2, target_purpose="Retail bank support bot",)print(ra.overview.to_df())ra.save(to="./results/")
# To run with a standards preset (cannot be combined with vulnerabilities)# from deepteam.frameworks import OWASPTop10# red_team(model_callback=model_callback, framework=OWASPTop10(categories=["LLM_01"]))There are four building blocks181920.
model_callback: the hook into the target app. The returnedRTTurncan also carryretrieval_contextandtools_called, so RAG and tool-call context can be evaluated too- Vulnerability: each subtype has its own dedicated 0/1 judge
- Attack: single-turn attacks encode a base attack or rewrite it with an LLM. Multi-turn attacks refine themselves while watching the target’s responses
- RiskAssessment: returns pass rates per vulnerability and per attack method, plus each case’s input/output and judging rationale. Since 1.0.9 it also includes a CVSS-like score
On top of that there’s deepteam run config.yaml to run the same thing from YAML, deepteam scan --diff main..HEAD -f sarif (1.0.7+) to statically scan source code with an LLM and use it as a CI gate21, and seven guardrails22.
Coverage
Counting from the 1.0.9 source code, there are 36 built-in vulnerabilities + 1 custom (124 subtypes total), and 22 single-turn + 5 multi-turn attacks. That’s counted differently from the README’s “50+ vulnerabilities” and “20+ attacks”162319.
| Category | Main vulnerabilities |
|---|---|
| Data / privacy | PII Leakage, Prompt Leakage (secrets, system prompt) |
| Responsible AI | Bias, Toxicity, Child Protection, Ethics, Fairness |
| Security | BOLA, BFLA, RBAC, Shell Injection, SQL Injection, SSRF, Tool Metadata Poisoning, Cross-Context Retrieval |
| Business | Misinformation, Intellectual Property, Competition, Hallucination |
| Agents | Goal Theft, Recursive Hijacking, Excessive Agency, Indirect Instruction (injection via RAG / tool output), Tool Orchestration Abuse, Insecure Inter-Agent Communication, etc. |
Attack methods include encoding attacks like Base64 / ROT13 / Leetspeak, LLM-rewrite attacks like Roleplay / Multilingual / Emotional Manipulation, agent-oriented ones like Permission Escalation / Goal Redirection, and multi-turn ones like Linear / Tree / Crescendo Jailbreaking.
There are six standards presets24. Having just learned the OWASP LLM Top 10 from the book, it’s nice that it maps directly onto a test configuration.
| Class | Contents |
|---|---|
OWASPTop10 | OWASP Top 10 for LLM 2025 (LLM01–10) |
OWASP_ASI_2026 | OWASP Top 10 for Agentic Applications 2026 (ASI01–10) |
NIST | Only the Measure function (M.1–M.4) of the NIST AI RMF |
MITRE | Six tactics from MITRE ATLAS |
EUAIAct | EU AI Act Article 5 (prohibited practices) and Annex III (high-risk areas). Added in 1.0.7 but not mentioned in the README |
Aegis / BeaverTails | Use prompts from real datasets (not synthetic attacks) |
Caveat: the docs and the code disagree
This was what bothered me most about DeepTeam.
- In the code,
red_team()defaults togpt-4o-minifor both simulator and evaluator, withignore_errors=True. The docs, however, saygpt-3.5-turbo-0125/gpt-4oandFalse - The
MITREATLASclass in the docs’ example doesn’t exist in the code (it’s actuallyMITRE) - The EU AI Act preset isn’t in the README, so other comparison articles written from the README list it as “not available”
Checking actual behavior in the source is the safe bet, and when something conflicts with the book, looking at the source first seems wise too.
Other caveats:
- Text only. No multimodal attacks using images or audio
- If you use an overly capable model as the simulator, that model’s own safety filters kick in and make it harder to generate attacks (an official note)18
- Cost is roughly “number of vulnerability types × number of attacks × (attack generation + target call + judging)”. With
run_all_attacks=True, every combination runs and the count jumps - Telemetry (Sentry and PostHog) is on by default. Set both
DEEPTEAM_TELEMETRY_OPT_OUT=YESandDEEPEVAL_TELEMETRY_OPT_OUT=YES25 - All guardrails are LLM-as-judge, with an LLM call per guard. Watch production latency and cost22
What else is there? Quality evaluation alternatives
Now for the exploration part. First, quality evaluation tools that can stand in for DeepEval.
| Tool | Form / license | Stars | Focus | Online eval in production | Self-hosting / price |
|---|---|---|---|---|---|
| DeepEval1 | OSS library, Apache-2.0 | 18.6k | pytest-style unit tests | Via paid SaaS | Library free, SaaS $0–$2,000/mo |
| Ragas26 | OSS library, Apache-2.0 | 15.9k | RAG metrics and test set generation | No | Free. Development stalled |
| promptfoo27 | OSS CLI, MIT (owned by OpenAI) | 25.7k | Prompt/model regression tests, model comparison, red teaming | Enterprise | Free, Enterprise custom |
| Inspect AI28 | OSS, MIT (UK AISI) | 2.9k | Model and agent benchmarks | No | Free |
| LangSmith29 | SaaS (SDK is MIT) | — | Tracing + evaluation + experiments | Yes | $0, $39/seat. Self-hosting Enterprise only |
| Langfuse30 | OSS platform, MIT | 35.4k | Observability + evaluation | Yes | Self-host free, Cloud $0–$2,499/mo, Tokyo region available |
| Arize Phoenix31 | Source-available (ELv2) | 11.7k | OTel tracing + evaluation | Since 2026-07 | Self-host free |
| TruLens32 | OSS, MIT (Snowflake) | 3.6k | RAG Triad, agent evaluation | In-library | Free |
| MLflow GenAI33 | OSS platform, Apache-2.0 | 28.3k | Evaluation as part of the ML lifecycle | Yes | Free, Databricks paid |
| Opik34 | OSS platform, Apache-2.0 | 22.4k | Observability + evaluation + prompt optimization | Yes | Self-host free, Cloud $0 / $19+ |
| Braintrust35 | SaaS (SDK is MIT) | — | Experiment comparison UX, production logs | Yes | $0 / $249 |
| Evidently36 | OSS, Apache-2.0 | 8.0k | ML drift monitoring + LLM evaluation | When self-hosted | Free. Cloud discontinued |
The big three clouds also have managed evaluation services (AWS Bedrock / AgentCore Evaluations, Azure Foundry Evaluation, Google Gen AI Evaluation)373839. I left OpenAI Evals out of the table because its API shuts down on 2026-11-3040.
Here’s how the main ones relate to DeepEval.
- Ragas: the library that became the common vocabulary for RAG metrics; Langfuse, MLflow, Braintrust and others pull in Ragas metrics. But there have been no releases since 2026-01, and the repository moved26. Depending on it alone is risky
- promptfoo: a config-driven tool that runs “prompts × models × test cases” from YAML. It handles both evaluation and red teaming, making it the closest competitor to DeepEval + DeepTeam. Assumes Node.js27
- Inspect AI: built by the UK AI Security Institute. Runs agents in sandboxes and ships 200+ benchmarks2841. Good for reproducibly measuring model/agent capabilities and safety; not suited to app-level RAG quality evaluation
- Langfuse: the observability platform with the largest OSS community. Since 2025-06, all features — including LLM-as-judge, annotation queues, and experiments — are MIT, and a Tokyo region launched in 2026-043042. Self-hosting means running ClickHouse, Postgres, Redis, and S3
- MLflow GenAI: a hub that can run DeepEval, Ragas, Phoenix, and TruLens scorers together inside
mlflow.genai.evaluate()43 - Opik: the most feature-complete OSS platform under Apache-2.0. Paid Cloud plans start cheap at $19/month3444
- Braintrust: a SaaS with a polished experiment-comparison UI and CI integration that comments scores on PRs and gates merges45
Lined up like this, a division of labor emerges: DeepEval is strong at “unit tests during development”, while Langfuse, Opik, LangSmith, and Braintrust are strong at “production tracing and monitoring”. Since DeepEval’s OSS edition has no production monitoring, DeepEval and Langfuse / Opik are partners to combine rather than competitors.
What else is there? Red teaming alternatives
Next, alternatives to DeepTeam.
| Tool | Form / license | Stars | Multi-turn | Agents / MCP | Standards presets | Price |
|---|---|---|---|---|---|---|
| DeepTeam16 | OSS, Apache-2.0 | 3.0k | Yes (5 kinds) | Yes | OWASP LLM 2025 / Agentic 2026 / NIST / ATLAS / EU AI Act | Free |
| promptfoo46 | OSS, MIT (owned by OpenAI) | 25.7k | Yes (Crescendo, GOAT, Hydra, etc.) | Up to MCP and coding agents | All of the above + ISO 42001 / GDPR, etc. | Free (attack generation capped at 10k probes/month), Enterprise |
| PyRIT47 | OSS, MIT (Microsoft) | 4.6k | Yes (Crescendo, PAIR, TAP) | XPIA, A2A | Individual scorers only | Free |
| garak48 | OSS, Apache-2.0 (NVIDIA) | 9.4k | Yes (TAP, PAIR, GOAT) | agent_breaker | OWASP (2023 numbering), EU AI Act | Free, fully local |
| Giskard OSS v349 | OSS, Apache-2.0 | 5.9k | Yes | Python callable | OWASP LLM 2025 tags | Free, Hub is Enterprise |
| PurpleLlama / CyberSecEval 450 | OSS, MIT (Meta) | 4.4k | Limited | Cyberattack agents | MITRE ATT&CK | Free |
| Inspect + inspect_evals41 | OSS, MIT | 2.9k | Yes | AgentDojo, AgentThreatBench | Supports OWASP Agentic | Free |
How each differs from DeepTeam, in a line or two:
- promptfoo: ahead on coverage (157 plugins), standards support, and CI / web UI maturity46. However, its advanced attack generation uses promptfoo’s (now OpenAI’s) remote inference by default, and the free tier is capped at 10k probes per month51. If you don’t want data leaving your environment, want to stay in Python, or don’t want to lean on a particular vendor, DeepTeam becomes a candidate
- PyRIT: a toolkit for experts to hand-build attack campaigns. Strong at multimodal and advanced multi-turn attacks, but no standards presets and more code to write47. DeepTeam’s strength is that it “just runs”
- garak: the “Nmap for LLMs”, with a huge set of static probes, strong at auditing a model on its own48. For app-level evaluation including RAG context and tool calls, DeepTeam fits better
- Giskard v3: handles quality evaluation and security scanning in one, and can call DeepTeam and garak as scan generators49. It was just rewritten from scratch for agents in 2026-08
While searching, I found the industry being reshuffled everywhere
As I researched the alternatives one by one, I noticed that “which company owns this tool now?” kept changing. Lined up, it looks like this.
| Event | When | Impact |
|---|---|---|
| OpenAI acquires promptfoo52 | Announced 2026-03-09 | Stated it will keep the MIT license27. Vendor neutrality is the thing to watch |
| OpenAI Evals (dashboard and API) deprecated40 | Read-only 2026-10-31, shut down 2026-11-30 | The official migration target is promptfoo53 |
| ClickHouse acquires Langfuse54 | 2026-01-16 | Stated no plans to change the license |
| Dynatrace acquires Arize AI55 | Agreement 2026-08-13 | The press release says nothing about Phoenix |
| Cisco acquires Galileo56 | Announced 2026-04 | Product renamed Splunk Agent Observability |
| W&B becomes part of CoreWeave57 | Completed 2025-05 | CoreWeave Agent Lens is pointed to as the successor to Weave |
| Evidently Cloud (SaaS) discontinued36 | Timing unconfirmed | OSS and self-hosting continue |
| Ragas repository moved26 | — | Last release 0.4.3 (2026-01-13) |
| Check Point acquires Lakera58, Palo Alto acquires Protect AI59 | 2025 | Guardrail and red-team products absorbed into big security vendors |
| LLM Guard (Protect AI) archived60 | — | Not for new adoption |
| Guardrails AI team joins Harvey61 | Announced 2026-09-09 | Future of the OSS unclear. Hosted inference ended 2026-08-25 |
| Giskard OSS v3 released49 | 2026-08-26 | Rewritten from scratch for agents. v2 out of maintenance |
Within a single year, in every area — quality evaluation, observability, red teaming, guardrails — major players were either acquired or wound down. Three things stood out to me.
First, OpenAI is shutting down its own evaluation service (OpenAI Evals) and pointing users to promptfoo, which it acquired, as the migration path4053. A company that builds models now owns a tool for evaluating models. promptfoo is a tool that lets you compare non-OpenAI models side by side, so I want to keep an eye on what happens to its neutrality.
Second, OSS observability is being absorbed by big data-infrastructure and monitoring players. Langfuse went to ClickHouse (the company behind the database Langfuse itself uses as its backend), Arize to Dynatrace, Galileo to Cisco (Splunk). LLM tracing is becoming one feature of existing monitoring stacks.
Third, turnover on the defensive side (guardrails). Here’s what the tools that plug holes found by red teaming in production look like:
| Tool | Status |
|---|---|
| NVIDIA NeMo Guardrails62 | Active (0.24.1, 2026-09). Docs assume pairing with garak |
| DeepTeam Guardrails22 | Seven LLM-judged guards. Easy, but each guard adds an LLM call |
| Meta LlamaFirewall50 | Last PyPI release 2025-05. Continuity unclear |
| Guardrails AI61 | Team moved to Harvey; hosted inference ended |
| LLM Guard60 | Archived |
It felt like stepping outside the tools I’d learned from the book only to find the ground itself moving. This table will probably look different again in six months, so please read the state described here as a snapshot from 2026-10.
Recommendations by use case
Summarizing the research’s recommendations:
| What you want to do | First choice | Notes / alternatives |
|---|---|---|
| Put evaluation of a Python app into CI | DeepEval | promptfoo if you want YAML and multiple models side by side |
| RAG quality evaluation | DeepEval’s RAG metrics | Ragas (watch the stall), TruLens, or MLflow to combine several libraries |
| Agent trajectory evaluation | DeepEval (tracing + Task Completion, etc.) | LangSmith for LangGraph, AgentCore Evaluations on AWS |
| Production observability + online evaluation | Langfuse (self-hosted or Tokyo region) | Opik (cheap, Apache-2.0), LangSmith (if you use LangChain) |
| Red teaming apps / agents (Python) | DeepTeam | Add garak for broad static-probe coverage |
| Red teaming with standards-compliance reports | promptfoo | DeepTeam if OpenAI ownership or remote generation concerns you |
| Expert-led, multimodal attack testing | PyRIT | AI Red Teaming Agent on Azure63 |
| Benchmarks for model selection | Inspect AI + inspect_evals | For Japanese: llm-jp-eval, Nejumi Leaderboard, etc.64 |
| Production guardrails | NeMo Guardrails | DeepTeam Guardrails for something lightweight |
Building an OSS-only stack around DeepEval and DeepTeam looks like this:
Development : DeepEval (regression tests on pytest, run on every PR)Pre-release : DeepTeam (red_team with OWASP LLM / Agentic presets, deepteam scan on the diff) + garak (static probes per model / endpoint)Production : Langfuse (tracing, online evaluation, annotation) + NeMo Guardrails or DeepTeam GuardrailsFeedback : Build Goldens from production traces and feed them back into DeepEval datasetsPaying for Confident AI lets you see development, production, and red teaming in one place, but even Starter is $200 a month, and the red-teaming module is Enterprise++11. If you build with OSS, handing production to Langfuse or Opik is the cost-sensible choice.
Notes for non-English (Japanese) use
I write from Japan, so this section is about Japanese, but most of it applies to any non-English app.
- Judge prompts mostly assume English. Every tool runs in Japanese, but judging quality isn’t officially validated. Ragas can adapt its prompts to the target language with
adapt()65; with DeepEval you can write G-Eval criteria in the target language, or pin down the judging steps with DAG. Either way, check agreement between the judge and a few dozen to a few hundred human labels in your language before setting thresholds - Cloud safety evaluation targets English. Azure Foundry’s safety evaluators aren’t available in Japanese regions, and English is stated as optimal66
- If you need data to stay in Japan, options are Langfuse Cloud Japan (Tokyo)42, AWS AgentCore Evaluations (Tokyo)38, Azure evaluation (Japan East / West, excluding safety evaluators)66, and Confident AI Enterprise11. LangSmith’s APAC region is Sydney; there’s no Japan region67
- Red-team attack prompts are English-centric. DeepTeam’s Multilingual attack switches languages to slip past defenses; it is not an attack set for Japanese-language apps. For those, you need to add CustomVulnerability or your own attack prompts
A local LLM as the judge might cut API costs
From here on, it’s my own idea, outside the research.
DeepEval’s top weakness was judge cost. QAG-style metrics call the judge many times per case, so cost grows as “test cases × metrics × calls per metric”. If you run it in CI on every PR, that multiplication runs every time.
Meanwhile, local LLMs have come a long way. DeepEval can use local OpenAI-compatible servers like Ollama or vLLM as the judge, and DeepTeam reuses DeepEval’s model layer. So for a company that already owns hardware like a DGX Spark, using an open frontier model that barely fits as the judge might cut API costs.
I have three reasons for thinking so.
- The workload shape fits. A judge’s job is to read fairly long retrieval context and answers and return a verdict with a short reason. In my post on putting Qwen3.8 Flash Next on two DGX Sparks into real work (Japanese), I wrote that jobs like code review — “read a long prompt, return something short” — feel fast even locally. What made the difference was TTFT, prefill, and aggregate throughput under parallel load, and evaluation fired in parallel with
deepeval test run -nuses exactly those - You can pin the judge. The weakness of API judges was scores shifting when the model is updated. Local weights don’t change unless you update them, so as a yardstick for release decisions, it might actually be more stable
- Data doesn’t have to leave. Evaluation data tends to include production inputs and outputs, and there are probably cases where evaluation never happened because the data couldn’t leave the company. That’s the same pattern as the “work we couldn’t do before because we couldn’t send it out” I described in the post above
There are preconditions, though. As in my post calculating the economics of two DGX Sparks (Japanese), this isn’t about buying new hardware to save on API fees. It works when the hardware is already there and you can run evaluations in its idle time. Also, if you switch the judge to an open model, calibration — checking agreement with human labels — is mandatory. Then again, that step is needed anyway with API judges if you’re evaluating in a non-English language, so it’s less extra work than a change in what the work you’d do anyway is applied to.
Adoption checklist
- Choice of judge model and cost estimate (test cases × metrics × calls per metric)
- Judge calibration (agreement with human labels), and whether to pin the judge model
- Telemetry opt-out (
DEEPEVAL_TELEMETRY_OPT_OUT,DEEPTEAM_TELEMETRY_OPT_OUT,RAGAS_DO_NOT_TRACK, etc.) - Version pinning. DeepEval has many breaking changes, so pin it in CI, e.g.
deepeval==4.2.x - For DeepTeam, check defaults in the code, not the docs
- Acquisition / shutdown status of the tools you use (promptfoo, Langfuse, Phoenix, Weave, OpenAI Evals)
Wrapping up
I learned about DeepEval and DeepTeam from a book, dug into them, and then went exploring.
Digging in, I found that DeepEval is still the first choice for putting evaluation into CI in Python, with the broadest metric range in open source. DeepTeam is a reasonable, easy red-teaming option for DeepEval users, with OWASP and EU AI Act presets ready to use. On the other hand, neither has production monitoring, so if you include production, pairing them with Langfuse or Opik is the realistic choice.
Exploring, I found that this field was reshuffled hard in 2026. promptfoo to OpenAI, Langfuse to ClickHouse, Arize to Dynatrace, Galileo to Cisco. OpenAI Evals stops at the end of next month. When choosing an evaluation tool, beyond the feature comparison table, be sure to check “who owns it now, and does it look likely to stay vendor-neutral?” One more thing: for judge costs, companies that own hardware like a DGX Spark might be able to hand the job to a local open model, and that’s something I’d like to try. DeepEval / DeepTeam are no exception to the reshuffle, either — nobody knows what will happen to Confident AI.
An introductory book gave me a map, and when I stepped outside, the terrain was changing every month. Better not to throw the map away — just check where you are, often.
References
- DeepEval GitHub https://github.com/confident-ai/deepeval
- DeepEval PyPI https://pypi.org/project/deepeval/
- DeepEval Releases https://github.com/confident-ai/deepeval/releases
- DeepEval Test Cases https://deepeval.com/docs/evaluation-test-cases
- DeepEval LLM Tracing https://deepeval.com/docs/evaluation-llm-tracing
- DeepEval Unit Testing in CI/CD https://deepeval.com/docs/evaluation-unit-testing-in-ci-cd
- DeepEval Metrics Introduction https://deepeval.com/docs/metrics-introduction
- DeepEval Classifiers https://deepeval.com/docs/classifiers-introduction
- DeepEval Synthetic Data Generation https://deepeval.com/docs/synthetic-data-generation-introduction
- DeepEval Integrations https://deepeval.com/integrations
- Confident AI Pricing https://www.confident-ai.com/pricing
- DeepEval Data Privacy https://deepeval.com/docs/data-privacy
- AgentCore Third-party Evaluators https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/third-party-evaluators.html
- Rubin Lake Technology Radar: DeepEval https://www.rubinlake.com/en/technology-radar/data-platforms-and-mlops/deepeval
- PromptLayer: DeepEval Review https://www.promptlayer.com/blog/deepeval-review-what-it-is-and-the-best-alternatives-in-2026/
- DeepTeam GitHub https://github.com/confident-ai/deepteam
- DeepTeam PyPI https://pypi.org/project/deepteam/
- DeepTeam Red Teaming Introduction https://www.trydeepteam.com/docs/red-teaming-introduction
- DeepTeam Adversarial Attacks https://www.trydeepteam.com/docs/red-teaming-adversarial-attacks
- DeepTeam Risk Assessment https://www.trydeepteam.com/docs/red-teaming-risk-assessment
- DeepTeam Code Scanning https://www.trydeepteam.com/docs/code-scanning-introduction
- DeepTeam Guardrails https://www.trydeepteam.com/docs/guardrails-introduction
- DeepTeam Vulnerabilities https://www.trydeepteam.com/docs/red-teaming-vulnerabilities
- DeepTeam Frameworks https://www.trydeepteam.com/docs/frameworks-introduction
- DeepTeam Data Privacy https://www.trydeepteam.com/docs/data-privacy
- Ragas GitHub https://github.com/vibrantlabsai/ragas
- promptfoo GitHub https://github.com/promptfoo/promptfoo
- Inspect AI GitHub https://github.com/UKGovernmentBEIS/inspect_ai
- LangSmith Pricing https://www.langchain.com/pricing
- Langfuse Pricing https://langfuse.com/pricing
- Arize Phoenix License https://arize.com/docs/phoenix/self-hosting/license
- TruLens GitHub https://github.com/truera/trulens
- MLflow GenAI Evaluation https://mlflow.org/docs/latest/genai/eval-monitor/
- Opik GitHub https://github.com/comet-ml/opik
- Braintrust Pricing https://www.braintrust.dev/pricing
- Evidently OSS vs Cloud https://docs.evidentlyai.com/faq/oss_vs_cloud
- Agent and model evaluations in Gemini Enterprise Agent Platform are now GA https://developers.googleblog.com/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/
- AgentCore Evaluations GA https://aws.amazon.com/about-aws/whats-new/2026/03/agentcore-evaluations-generally-available
- Microsoft Foundry Built-in Evaluators https://learn.microsoft.com/en-us/azure/foundry/concepts/built-in-evaluators
- OpenAI API Deprecations https://developers.openai.com/api/docs/deprecations
- inspect_evals GitHub https://github.com/UKGovernmentBEIS/inspect_evals
- Langfuse Cloud Japan https://langfuse.com/blog/2026-04-27-langfuse-cloud-japan
- MLflow Third-party Scorers https://mlflow.org/blog/third-party-scorers/
- Comet Pricing https://www.comet.com/site/pricing/
- Braintrust Run in CI https://braintrust.dev/docs/evaluate/run-in-ci
- promptfoo Red Team Plugins https://www.promptfoo.dev/docs/red-team/plugins/
- PyRIT GitHub https://github.com/microsoft/PyRIT
- garak GitHub https://github.com/NVIDIA/garak
- Giskard OSS GitHub https://github.com/Giskard-AI/giskard-oss
- PurpleLlama GitHub https://github.com/meta-llama/PurpleLlama
- promptfoo Pricing https://www.promptfoo.dev/pricing/
- OpenAI to acquire Promptfoo https://openai.com/index/openai-to-acquire-promptfoo/
- Moving from OpenAI Evals to promptfoo https://developers.openai.com/cookbook/examples/evaluation/moving-from-openai-evals-to-promptfoo
- Langfuse joins ClickHouse https://langfuse.com/blog/announcing-acquisition
- Dynatrace to acquire Arize https://www.dynatrace.com/news/press-release/dynatrace-to-acquire-arize/
- Splunk / Galileo https://www.splunk.com/en_us/about-splunk/acquisitions/galileo.html
- CoreWeave completes acquisition of Weights & Biases https://coreweave.com/blog/coreweave-completes-acquisition-of-weights-biases
- Check Point acquires Lakera https://www.checkpoint.com/press-releases/check-point-acquires-lakera/
- Palo Alto Networks completes acquisition of Protect AI https://www.paloaltonetworks.com/company/press/2025/palo-alto-networks-completes-acquisition-of-protect-ai
- LLM Guard GitHub https://github.com/protectai/llm-guard
- Guardrails AI joins Harvey https://www.harvey.ai/blog/guardrails-ai-joins-harvey
- NeMo Guardrails GitHub https://github.com/NVIDIA-NeMo/Guardrails
- Azure AI Red Teaming Agent https://learn.microsoft.com/azure/ai-foundry/concepts/ai-red-teaming-agent
- Nejumi LLM Leaderboard https://nejumi.ai
- Ragas Metrics Language Adaptation https://docs.ragas.io/en/stable/howtos/customizations/metrics/metrics_language_adaptation/
- Microsoft Foundry Evaluation Regions and Limits https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-regions-limits-virtual-network
- LangSmith Regions FAQ https://docs.langchain.com/langsmith/regions-faq