4666 文字
23 分
What else is there besides DeepEval? Looking for LLM app evaluation alternatives, I found 2026 was full of acquisitions (promptfoo went to OpenAI, Langfuse to ClickHouse)
About this post

Co-edited with Claude (Anthropic).

Summary
  • Evaluating GenAI apps splits into two tracks: quality evaluation, which checks accuracy and RAG faithfulness, and red teaming, which hunts for jailbreaks and data leaks. DeepEval covers the former, DeepTeam the latter
  • DeepEval is an evaluation library that runs on pytest, with 50+ metrics. It covers RAG, agents, conversations, MCP, and even voice — the broadest range in open source. Its weak points are the cost and score variance of LLM-as-judge, plus frequent breaking changes (4.2.0 flipped the direction of some scores)
  • DeepTeam is a red-teaming tool that hooks into your app through a single Python callback, with presets for OWASP LLM Top 10, the EU AI Act, and more. But its docs and code disagree in noticeable ways — even the default models differ
  • Alternatives: for quality evaluation, promptfoo, Langfuse, Opik, Ragas, Inspect AI and others; for red teaming, promptfoo, PyRIT, garak, Giskard and others. DeepEval / DeepTeam have no production monitoring, so pairing them with Langfuse or Opik for production is the realistic setup
  • What surprised me most while looking for alternatives was how hard the industry is being reshuffled. promptfoo was acquired by OpenAI, Langfuse by ClickHouse, Arize is being acquired by Dynatrace, Galileo went to Cisco, and OpenAI Evals shuts down on 2026-11-30. When choosing a tool, check not just the features but “who owns it now, and will it stay vendor-neutral?”
  • As an idea outside the research: judge costs might be cut with a local LLM. A company that already owns hardware like a DGX Spark could use an open frontier model as the judge, saving API spend while also pinning the judge and keeping data in-house. Calibration is still mandatory

Introduction#

I read 『生成AIアプリケーション評価入門』 (roughly “Introduction to Evaluating Generative AI Applications”, Shinsuke Matsuki, Gijutsu-Hyoronsha) — a Japanese book. It starts from classic evaluation metrics like the confusion matrix and BLEU / ROUGE, moves through the ISO/IEC 25059 quality model and the OWASP LLM Top 10, and finishes with a hands-on part that uses DeepEval and DeepTeam for evaluation and red teaming. As a systematic introduction it was easy to read and not too thick, which made it just right for getting the big picture of evaluation.

* Contains Amazon Associates links

After finishing the book, two things bugged me. First, how far can DeepEval and DeepTeam actually go, and where are their weak spots? Second, are there other options? DeepEval can’t be the only evaluation tool out there, and without comparing, I can’t tell where DeepEval stands.

So I dug into DeepEval and DeepTeam through their official docs and the source code of the published packages, and then went exploring for alternatives. As I searched, something caught my eye before any feature differences did: this industry was being reshuffled at a furious pace in 2026. Roughly half of the major tools were acquired, shut down their service, or stalled in development within the past year.

Note that I haven’t run DeepEval / DeepTeam myself. This is a summary of what I learned from the book plus the official docs, PyPI, GitHub, and source code. Versions, star counts, and prices are all as of 2026-10-05.

Evaluation splits into two tracks#

To set things up first: “evaluation” of GenAI apps broadly splits into two. The book also covers DeepEval and DeepTeam in separate chapters.

TrackWhat it checksCovered byMain competitors
Quality evaluation (Eval)Accuracy, faithfulness, RAG retrieval quality, agent task completion, etc.DeepEvalRagas, promptfoo, Langfuse, LangSmith, Opik, MLflow, Braintrust, Inspect AI, cloud evaluation services
Safety evaluation (Red Teaming)Jailbreaks, prompt injection, data leakage, excessive agency, etc.DeepTeampromptfoo, PyRIT, garak, Giskard, commercial (Prisma AIRS, Lakera)

Both are open source from the same company, Confident AI, and DeepTeam depends on DeepEval at runtime. Below, I dig into both first, then look at alternatives for each.

Digging into DeepEval#

Overview#

ItemDetails
DeveloperConfident AI
Positioning”Pytest for LLM apps”. Metrics run locally1
LicenseApache-2.02
LatestPython 4.2.8 (2026-10-02). Python 3.9+2. A TypeScript SDK also exists
GitHub~18.6k stars, releases almost weekly13

Version 4.0 (2026-05) added an evaluation harness for coding agents such as Claude Code / Cursor / Codex3. You can see from the tooling side, too, that the evaluation target is shifting from “chatbots” to “agents”.

Basic usage#

There are only a few core concepts45.

  • LLMTestCase: a single input/output exchange. Has input, actual_output, expected_output, retrieval_context, tools_called, etc.
  • ConversationalTestCase: a multi-turn conversation. ConversationSimulator can generate conversations automatically
  • Golden / EvaluationDataset: test data that holds only the expectations. The app is run at evaluation time to fill in the outputs
  • evaluate(): the entry point for evaluating from a script
  • assert_test() and deepeval test run: the entry point for running on pytest in CI. Has flags for parallel runs (-n), caching (-c), and repeats (-r)6
  • @observe(): captures traces and applies metrics per span or to the whole trace

Following the official style, a test looks like this (as of 4.2.x).

# test_app.py → `deepeval test run test_app.py`
from deepeval import assert_test
from deepeval.metrics import GEval, AnswerRelevancyMetric, FaithfulnessMetric
from deepeval.test_case import LLMTestCase, SingleTurnParams
correctness = GEval(
name="Correctness",
criteria="Determine if the 'actual output' is correct based on the 'expected output'.",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.EXPECTED_OUTPUT],
threshold=0.5,
)
def test_refund():
tc = LLMTestCase(
input="What if these shoes don't fit?",
actual_output="You have 30 days to get a full refund at no extra cost.",
expected_output="We offer a 30-day full refund at no extra costs.",
retrieval_context=["All customers are eligible for a 30 day full refund at no extra costs."],
)
assert_test(tc, [correctness, AnswerRelevancyMetric(threshold=0.7), FaithfulnessMetric()])

To evaluate an agent, combine it with tracing.

from deepeval.tracing import observe
from deepeval.dataset import EvaluationDataset, Golden
from deepeval.metrics import TaskCompletionMetric
@observe()
def agent(q): ...
ds = EvaluationDataset(goldens=[Golden(input="Plan a two-day trip to Paris")])
for g in ds.evals_iterator(metrics=[TaskCompletionMetric()]):
agent(g.input)

DeepEval’s defining trait is that evaluations look just like ordinary unit tests, which lowers the psychological barrier to putting them in CI.

Metrics#

The official docs advertise “50+“7. Here are the main ones by category.

CategoryMetrics
General / customG-Eval (criteria written in natural language), DAG (decision tree that pins down the judging steps), Arena G-Eval (pairwise comparison), custom metrics
RAGAnswer Relevancy, Faithfulness, Contextual Precision / Recall / Relevancy
AgentsTask Completion, Step Efficiency, Plan Adherence, Plan Quality (whole trajectory) / Tool Correctness, Argument Correctness (per component)
ConversationKnowledge Retention, Role Adherence, Conversation Completeness, Turn Relevancy, etc.
SafetyBias, Toxicity, PII Leakage, Misuse, Non-Advice, Role Violation
MCPMCP Task Completion, MCP Use, Multi-Turn MCP Use
MultimodalImage generation / editing evaluation, metrics for voice agents (4.1.10+)
Non-LLMJSON Correctness, Exact Match, Pattern Match
Classifiers (4.2+)Judges that return labels instead of scores. Refusal, Escalation, Prompt Injection, Data Leakage, etc.8

Things to know about how it works:

  • Almost everything is LLM-as-judge. Each metric uses one of QAG (extract claims and judge them one by one), G-Eval, or DAG
  • Scores range from 0 to 1 and come with a reason. The default threshold is 0.5; strict_mode=True makes it a 0/1 judgment
  • In 4.2.0 every metric was unified to “higher is better”. Before that, Bias and Toxicity pointed the other way, so thresholds in old code need revisiting
  • The default judge is OpenAI (gpt-5.4 in 4.2.8). Anthropic, Bedrock, Gemini, local OpenAI-compatible servers like Ollama and vLLM, LiteLLM, and more are built in, and you switch with deepeval set-<provider>1

There’s also synthetic data generation that builds evaluation data (Goldens) from documents, benchmarks like MMLU and GSM8K, prompt optimization via GEPA / MIPROv2, and integrations with LangChain / LlamaIndex / OpenAI Agents SDK / Anthropic SDK and others910. On synthetic data, I found it honest that the official docs themselves recommend the order “human-curated data first, then production traffic, synthetic data last”9.

Data handling and the paid cloud#

DeepEval itself is free, but team dashboards, online evaluation, and annotation live on the paid Confident AI side. Prices as of 2026-10-0511:

PlanMonthlyMain contents
Free$02 seats, 1 project, 5 test runs per week
Starter$200Unlimited seats, 5 projects, online evaluation, annotation queues
Team$2,000Unlimited projects, versioning, RBAC, SSO
EnterpriseCustom quoteOn-prem, data residency options (including Japan). The red-teaming module is “Enterprise++”

It used to start at $19.99 per user per month, so the pricing structure has changed a lot.

Three things to watch in data handling12:

  • Evaluation data is not sent to the cloud unless you run deepeval login or set CONFIDENT_API_KEY
  • However, anonymous telemetry is on by default (sent to PostHog). Stop it with DEEPEVAL_TELEMETRY_OPT_OUT=1
  • It auto-loads .env / .env.local on import. In CI, DEEPEVAL_DISABLE_DOTENV=1 is recommended

Strengths and weaknesses#

Strengths:

  • The broadest metric set in open source, all the way to agents, MCP, and voice
  • Evaluations are written as code and go straight into CI via pytest
  • Many judge model options, including fully local setups with Ollama or vLLM
  • AWS’s Bedrock AgentCore Evaluations adopted DeepEval metrics as third-party evaluators13

Weaknesses:

  • Cost and latency. QAG-style metrics call the judge many times per case, so as test counts grow, token costs and rate limits start to bite14
  • Scores fluctuate. Scores shift when the judge model is updated, so you need to check against human labels before using them for release decisions14
  • Lots of API changes. The rename from LLMTestCaseParams to SingleTurnParams, the score-direction flip in 4.2.0, the DAG overhaul, and so on. Old tutorials often don’t run
  • Only engineers can write evaluations. Team sharing, history, and production monitoring are on the paid side15

As someone who learned from a book, it’s worth remembering that code from the book won’t necessarily run as-is on the current version.

Digging into DeepTeam#

Overview#

ItemDetails
Positioning”Penetration testing for LLMs”. Auto-generates attacks to surface app vulnerabilities16
LicenseApache-2.017
Latest1.0.9 (PyPI, 2026-08-12). Python 3.9 to below 3.1417
GitHub~3.0k stars. Created 2025-03; first stable 1.0.0 on 2025-11-121617
Relation to DeepEvalA separate package, but depends on DeepEval at runtime and reuses its model layer and test case types18

GitHub Releases has only three entries, and the v1.0.9 tag carries the 1.0.0 release notes, so treat PyPI as the source of truth for versions.

Basic usage#

from deepteam import red_team
from deepteam.vulnerabilities import Bias, PIILeakage
from deepteam.attacks.single_turn import PromptInjection, ROT13
from deepteam.attacks.multi_turn import LinearJailbreaking
from deepteam.test_case import RTTurn
async def model_callback(input: str, turns: list[RTTurn] | None = None):
reply = await my_app.generate(input) # the app under test
return RTTurn(role="assistant", content=reply) # returning a plain str also works
ra = red_team(
model_callback=model_callback,
vulnerabilities=[Bias(types=["race", "gender"]), PIILeakage(types=["direct_disclosure"])],
attacks=[PromptInjection(weight=2), ROT13(), LinearJailbreaking(num_turns=3)],
simulator_model="gpt-4o-mini", evaluation_model="gpt-4o",
attacks_per_vulnerability_type=2,
target_purpose="Retail bank support bot",
)
print(ra.overview.to_df())
ra.save(to="./results/")
# To run with a standards preset (cannot be combined with vulnerabilities)
# from deepteam.frameworks import OWASPTop10
# red_team(model_callback=model_callback, framework=OWASPTop10(categories=["LLM_01"]))

There are four building blocks181920.

  • model_callback: the hook into the target app. The returned RTTurn can also carry retrieval_context and tools_called, so RAG and tool-call context can be evaluated too
  • Vulnerability: each subtype has its own dedicated 0/1 judge
  • Attack: single-turn attacks encode a base attack or rewrite it with an LLM. Multi-turn attacks refine themselves while watching the target’s responses
  • RiskAssessment: returns pass rates per vulnerability and per attack method, plus each case’s input/output and judging rationale. Since 1.0.9 it also includes a CVSS-like score

On top of that there’s deepteam run config.yaml to run the same thing from YAML, deepteam scan --diff main..HEAD -f sarif (1.0.7+) to statically scan source code with an LLM and use it as a CI gate21, and seven guardrails22.

Coverage#

Counting from the 1.0.9 source code, there are 36 built-in vulnerabilities + 1 custom (124 subtypes total), and 22 single-turn + 5 multi-turn attacks. That’s counted differently from the README’s “50+ vulnerabilities” and “20+ attacks”162319.

CategoryMain vulnerabilities
Data / privacyPII Leakage, Prompt Leakage (secrets, system prompt)
Responsible AIBias, Toxicity, Child Protection, Ethics, Fairness
SecurityBOLA, BFLA, RBAC, Shell Injection, SQL Injection, SSRF, Tool Metadata Poisoning, Cross-Context Retrieval
BusinessMisinformation, Intellectual Property, Competition, Hallucination
AgentsGoal Theft, Recursive Hijacking, Excessive Agency, Indirect Instruction (injection via RAG / tool output), Tool Orchestration Abuse, Insecure Inter-Agent Communication, etc.

Attack methods include encoding attacks like Base64 / ROT13 / Leetspeak, LLM-rewrite attacks like Roleplay / Multilingual / Emotional Manipulation, agent-oriented ones like Permission Escalation / Goal Redirection, and multi-turn ones like Linear / Tree / Crescendo Jailbreaking.

There are six standards presets24. Having just learned the OWASP LLM Top 10 from the book, it’s nice that it maps directly onto a test configuration.

ClassContents
OWASPTop10OWASP Top 10 for LLM 2025 (LLM01–10)
OWASP_ASI_2026OWASP Top 10 for Agentic Applications 2026 (ASI01–10)
NISTOnly the Measure function (M.1–M.4) of the NIST AI RMF
MITRESix tactics from MITRE ATLAS
EUAIActEU AI Act Article 5 (prohibited practices) and Annex III (high-risk areas). Added in 1.0.7 but not mentioned in the README
Aegis / BeaverTailsUse prompts from real datasets (not synthetic attacks)

Caveat: the docs and the code disagree#

This was what bothered me most about DeepTeam.

  • In the code, red_team() defaults to gpt-4o-mini for both simulator and evaluator, with ignore_errors=True. The docs, however, say gpt-3.5-turbo-0125 / gpt-4o and False
  • The MITREATLAS class in the docs’ example doesn’t exist in the code (it’s actually MITRE)
  • The EU AI Act preset isn’t in the README, so other comparison articles written from the README list it as “not available”

Checking actual behavior in the source is the safe bet, and when something conflicts with the book, looking at the source first seems wise too.

Other caveats:

  • Text only. No multimodal attacks using images or audio
  • If you use an overly capable model as the simulator, that model’s own safety filters kick in and make it harder to generate attacks (an official note)18
  • Cost is roughly “number of vulnerability types × number of attacks × (attack generation + target call + judging)”. With run_all_attacks=True, every combination runs and the count jumps
  • Telemetry (Sentry and PostHog) is on by default. Set both DEEPTEAM_TELEMETRY_OPT_OUT=YES and DEEPEVAL_TELEMETRY_OPT_OUT=YES25
  • All guardrails are LLM-as-judge, with an LLM call per guard. Watch production latency and cost22

What else is there? Quality evaluation alternatives#

Now for the exploration part. First, quality evaluation tools that can stand in for DeepEval.

ToolForm / licenseStarsFocusOnline eval in productionSelf-hosting / price
DeepEval1OSS library, Apache-2.018.6kpytest-style unit testsVia paid SaaSLibrary free, SaaS $0–$2,000/mo
Ragas26OSS library, Apache-2.015.9kRAG metrics and test set generationNoFree. Development stalled
promptfoo27OSS CLI, MIT (owned by OpenAI)25.7kPrompt/model regression tests, model comparison, red teamingEnterpriseFree, Enterprise custom
Inspect AI28OSS, MIT (UK AISI)2.9kModel and agent benchmarksNoFree
LangSmith29SaaS (SDK is MIT)—Tracing + evaluation + experimentsYes$0, $39/seat. Self-hosting Enterprise only
Langfuse30OSS platform, MIT35.4kObservability + evaluationYesSelf-host free, Cloud $0–$2,499/mo, Tokyo region available
Arize Phoenix31Source-available (ELv2)11.7kOTel tracing + evaluationSince 2026-07Self-host free
TruLens32OSS, MIT (Snowflake)3.6kRAG Triad, agent evaluationIn-libraryFree
MLflow GenAI33OSS platform, Apache-2.028.3kEvaluation as part of the ML lifecycleYesFree, Databricks paid
Opik34OSS platform, Apache-2.022.4kObservability + evaluation + prompt optimizationYesSelf-host free, Cloud $0 / $19+
Braintrust35SaaS (SDK is MIT)—Experiment comparison UX, production logsYes$0 / $249
Evidently36OSS, Apache-2.08.0kML drift monitoring + LLM evaluationWhen self-hostedFree. Cloud discontinued

The big three clouds also have managed evaluation services (AWS Bedrock / AgentCore Evaluations, Azure Foundry Evaluation, Google Gen AI Evaluation)373839. I left OpenAI Evals out of the table because its API shuts down on 2026-11-3040.

Here’s how the main ones relate to DeepEval.

  • Ragas: the library that became the common vocabulary for RAG metrics; Langfuse, MLflow, Braintrust and others pull in Ragas metrics. But there have been no releases since 2026-01, and the repository moved26. Depending on it alone is risky
  • promptfoo: a config-driven tool that runs “prompts × models × test cases” from YAML. It handles both evaluation and red teaming, making it the closest competitor to DeepEval + DeepTeam. Assumes Node.js27
  • Inspect AI: built by the UK AI Security Institute. Runs agents in sandboxes and ships 200+ benchmarks2841. Good for reproducibly measuring model/agent capabilities and safety; not suited to app-level RAG quality evaluation
  • Langfuse: the observability platform with the largest OSS community. Since 2025-06, all features — including LLM-as-judge, annotation queues, and experiments — are MIT, and a Tokyo region launched in 2026-043042. Self-hosting means running ClickHouse, Postgres, Redis, and S3
  • MLflow GenAI: a hub that can run DeepEval, Ragas, Phoenix, and TruLens scorers together inside mlflow.genai.evaluate()43
  • Opik: the most feature-complete OSS platform under Apache-2.0. Paid Cloud plans start cheap at $19/month3444
  • Braintrust: a SaaS with a polished experiment-comparison UI and CI integration that comments scores on PRs and gates merges45

Lined up like this, a division of labor emerges: DeepEval is strong at “unit tests during development”, while Langfuse, Opik, LangSmith, and Braintrust are strong at “production tracing and monitoring”. Since DeepEval’s OSS edition has no production monitoring, DeepEval and Langfuse / Opik are partners to combine rather than competitors.

What else is there? Red teaming alternatives#

Next, alternatives to DeepTeam.

ToolForm / licenseStarsMulti-turnAgents / MCPStandards presetsPrice
DeepTeam16OSS, Apache-2.03.0kYes (5 kinds)YesOWASP LLM 2025 / Agentic 2026 / NIST / ATLAS / EU AI ActFree
promptfoo46OSS, MIT (owned by OpenAI)25.7kYes (Crescendo, GOAT, Hydra, etc.)Up to MCP and coding agentsAll of the above + ISO 42001 / GDPR, etc.Free (attack generation capped at 10k probes/month), Enterprise
PyRIT47OSS, MIT (Microsoft)4.6kYes (Crescendo, PAIR, TAP)XPIA, A2AIndividual scorers onlyFree
garak48OSS, Apache-2.0 (NVIDIA)9.4kYes (TAP, PAIR, GOAT)agent_breakerOWASP (2023 numbering), EU AI ActFree, fully local
Giskard OSS v349OSS, Apache-2.05.9kYesPython callableOWASP LLM 2025 tagsFree, Hub is Enterprise
PurpleLlama / CyberSecEval 450OSS, MIT (Meta)4.4kLimitedCyberattack agentsMITRE ATT&CKFree
Inspect + inspect_evals41OSS, MIT2.9kYesAgentDojo, AgentThreatBenchSupports OWASP AgenticFree

How each differs from DeepTeam, in a line or two:

  • promptfoo: ahead on coverage (157 plugins), standards support, and CI / web UI maturity46. However, its advanced attack generation uses promptfoo’s (now OpenAI’s) remote inference by default, and the free tier is capped at 10k probes per month51. If you don’t want data leaving your environment, want to stay in Python, or don’t want to lean on a particular vendor, DeepTeam becomes a candidate
  • PyRIT: a toolkit for experts to hand-build attack campaigns. Strong at multimodal and advanced multi-turn attacks, but no standards presets and more code to write47. DeepTeam’s strength is that it “just runs”
  • garak: the “Nmap for LLMs”, with a huge set of static probes, strong at auditing a model on its own48. For app-level evaluation including RAG context and tool calls, DeepTeam fits better
  • Giskard v3: handles quality evaluation and security scanning in one, and can call DeepTeam and garak as scan generators49. It was just rewritten from scratch for agents in 2026-08

While searching, I found the industry being reshuffled everywhere#

As I researched the alternatives one by one, I noticed that “which company owns this tool now?” kept changing. Lined up, it looks like this.

EventWhenImpact
OpenAI acquires promptfoo52Announced 2026-03-09Stated it will keep the MIT license27. Vendor neutrality is the thing to watch
OpenAI Evals (dashboard and API) deprecated40Read-only 2026-10-31, shut down 2026-11-30The official migration target is promptfoo53
ClickHouse acquires Langfuse542026-01-16Stated no plans to change the license
Dynatrace acquires Arize AI55Agreement 2026-08-13The press release says nothing about Phoenix
Cisco acquires Galileo56Announced 2026-04Product renamed Splunk Agent Observability
W&B becomes part of CoreWeave57Completed 2025-05CoreWeave Agent Lens is pointed to as the successor to Weave
Evidently Cloud (SaaS) discontinued36Timing unconfirmedOSS and self-hosting continue
Ragas repository moved26—Last release 0.4.3 (2026-01-13)
Check Point acquires Lakera58, Palo Alto acquires Protect AI592025Guardrail and red-team products absorbed into big security vendors
LLM Guard (Protect AI) archived60—Not for new adoption
Guardrails AI team joins Harvey61Announced 2026-09-09Future of the OSS unclear. Hosted inference ended 2026-08-25
Giskard OSS v3 released492026-08-26Rewritten from scratch for agents. v2 out of maintenance

Within a single year, in every area — quality evaluation, observability, red teaming, guardrails — major players were either acquired or wound down. Three things stood out to me.

First, OpenAI is shutting down its own evaluation service (OpenAI Evals) and pointing users to promptfoo, which it acquired, as the migration path4053. A company that builds models now owns a tool for evaluating models. promptfoo is a tool that lets you compare non-OpenAI models side by side, so I want to keep an eye on what happens to its neutrality.

Second, OSS observability is being absorbed by big data-infrastructure and monitoring players. Langfuse went to ClickHouse (the company behind the database Langfuse itself uses as its backend), Arize to Dynatrace, Galileo to Cisco (Splunk). LLM tracing is becoming one feature of existing monitoring stacks.

Third, turnover on the defensive side (guardrails). Here’s what the tools that plug holes found by red teaming in production look like:

ToolStatus
NVIDIA NeMo Guardrails62Active (0.24.1, 2026-09). Docs assume pairing with garak
DeepTeam Guardrails22Seven LLM-judged guards. Easy, but each guard adds an LLM call
Meta LlamaFirewall50Last PyPI release 2025-05. Continuity unclear
Guardrails AI61Team moved to Harvey; hosted inference ended
LLM Guard60Archived

It felt like stepping outside the tools I’d learned from the book only to find the ground itself moving. This table will probably look different again in six months, so please read the state described here as a snapshot from 2026-10.

Recommendations by use case#

Summarizing the research’s recommendations:

What you want to doFirst choiceNotes / alternatives
Put evaluation of a Python app into CIDeepEvalpromptfoo if you want YAML and multiple models side by side
RAG quality evaluationDeepEval’s RAG metricsRagas (watch the stall), TruLens, or MLflow to combine several libraries
Agent trajectory evaluationDeepEval (tracing + Task Completion, etc.)LangSmith for LangGraph, AgentCore Evaluations on AWS
Production observability + online evaluationLangfuse (self-hosted or Tokyo region)Opik (cheap, Apache-2.0), LangSmith (if you use LangChain)
Red teaming apps / agents (Python)DeepTeamAdd garak for broad static-probe coverage
Red teaming with standards-compliance reportspromptfooDeepTeam if OpenAI ownership or remote generation concerns you
Expert-led, multimodal attack testingPyRITAI Red Teaming Agent on Azure63
Benchmarks for model selectionInspect AI + inspect_evalsFor Japanese: llm-jp-eval, Nejumi Leaderboard, etc.64
Production guardrailsNeMo GuardrailsDeepTeam Guardrails for something lightweight

Building an OSS-only stack around DeepEval and DeepTeam looks like this:

Development : DeepEval (regression tests on pytest, run on every PR)
Pre-release : DeepTeam (red_team with OWASP LLM / Agentic presets, deepteam scan on the diff)
+ garak (static probes per model / endpoint)
Production : Langfuse (tracing, online evaluation, annotation)
+ NeMo Guardrails or DeepTeam Guardrails
Feedback : Build Goldens from production traces and feed them back into DeepEval datasets

Paying for Confident AI lets you see development, production, and red teaming in one place, but even Starter is $200 a month, and the red-teaming module is Enterprise++11. If you build with OSS, handing production to Langfuse or Opik is the cost-sensible choice.

Notes for non-English (Japanese) use#

I write from Japan, so this section is about Japanese, but most of it applies to any non-English app.

  • Judge prompts mostly assume English. Every tool runs in Japanese, but judging quality isn’t officially validated. Ragas can adapt its prompts to the target language with adapt()65; with DeepEval you can write G-Eval criteria in the target language, or pin down the judging steps with DAG. Either way, check agreement between the judge and a few dozen to a few hundred human labels in your language before setting thresholds
  • Cloud safety evaluation targets English. Azure Foundry’s safety evaluators aren’t available in Japanese regions, and English is stated as optimal66
  • If you need data to stay in Japan, options are Langfuse Cloud Japan (Tokyo)42, AWS AgentCore Evaluations (Tokyo)38, Azure evaluation (Japan East / West, excluding safety evaluators)66, and Confident AI Enterprise11. LangSmith’s APAC region is Sydney; there’s no Japan region67
  • Red-team attack prompts are English-centric. DeepTeam’s Multilingual attack switches languages to slip past defenses; it is not an attack set for Japanese-language apps. For those, you need to add CustomVulnerability or your own attack prompts

A local LLM as the judge might cut API costs#

From here on, it’s my own idea, outside the research.

DeepEval’s top weakness was judge cost. QAG-style metrics call the judge many times per case, so cost grows as “test cases × metrics × calls per metric”. If you run it in CI on every PR, that multiplication runs every time.

Meanwhile, local LLMs have come a long way. DeepEval can use local OpenAI-compatible servers like Ollama or vLLM as the judge, and DeepTeam reuses DeepEval’s model layer. So for a company that already owns hardware like a DGX Spark, using an open frontier model that barely fits as the judge might cut API costs.

I have three reasons for thinking so.

  • The workload shape fits. A judge’s job is to read fairly long retrieval context and answers and return a verdict with a short reason. In my post on putting Qwen3.8 Flash Next on two DGX Sparks into real work (Japanese), I wrote that jobs like code review — “read a long prompt, return something short” — feel fast even locally. What made the difference was TTFT, prefill, and aggregate throughput under parallel load, and evaluation fired in parallel with deepeval test run -n uses exactly those
  • You can pin the judge. The weakness of API judges was scores shifting when the model is updated. Local weights don’t change unless you update them, so as a yardstick for release decisions, it might actually be more stable
  • Data doesn’t have to leave. Evaluation data tends to include production inputs and outputs, and there are probably cases where evaluation never happened because the data couldn’t leave the company. That’s the same pattern as the “work we couldn’t do before because we couldn’t send it out” I described in the post above

There are preconditions, though. As in my post calculating the economics of two DGX Sparks (Japanese), this isn’t about buying new hardware to save on API fees. It works when the hardware is already there and you can run evaluations in its idle time. Also, if you switch the judge to an open model, calibration — checking agreement with human labels — is mandatory. Then again, that step is needed anyway with API judges if you’re evaluating in a non-English language, so it’s less extra work than a change in what the work you’d do anyway is applied to.

Adoption checklist#

  • Choice of judge model and cost estimate (test cases × metrics × calls per metric)
  • Judge calibration (agreement with human labels), and whether to pin the judge model
  • Telemetry opt-out (DEEPEVAL_TELEMETRY_OPT_OUT, DEEPTEAM_TELEMETRY_OPT_OUT, RAGAS_DO_NOT_TRACK, etc.)
  • Version pinning. DeepEval has many breaking changes, so pin it in CI, e.g. deepeval==4.2.x
  • For DeepTeam, check defaults in the code, not the docs
  • Acquisition / shutdown status of the tools you use (promptfoo, Langfuse, Phoenix, Weave, OpenAI Evals)

Wrapping up#

I learned about DeepEval and DeepTeam from a book, dug into them, and then went exploring.

Digging in, I found that DeepEval is still the first choice for putting evaluation into CI in Python, with the broadest metric range in open source. DeepTeam is a reasonable, easy red-teaming option for DeepEval users, with OWASP and EU AI Act presets ready to use. On the other hand, neither has production monitoring, so if you include production, pairing them with Langfuse or Opik is the realistic choice.

Exploring, I found that this field was reshuffled hard in 2026. promptfoo to OpenAI, Langfuse to ClickHouse, Arize to Dynatrace, Galileo to Cisco. OpenAI Evals stops at the end of next month. When choosing an evaluation tool, beyond the feature comparison table, be sure to check “who owns it now, and does it look likely to stay vendor-neutral?” One more thing: for judge costs, companies that own hardware like a DGX Spark might be able to hand the job to a local open model, and that’s something I’d like to try. DeepEval / DeepTeam are no exception to the reshuffle, either — nobody knows what will happen to Confident AI.

An introductory book gave me a map, and when I stepped outside, the terrain was changing every month. Better not to throw the map away — just check where you are, often.

References#

  1. DeepEval GitHub https://github.com/confident-ai/deepeval
  2. DeepEval PyPI https://pypi.org/project/deepeval/
  3. DeepEval Releases https://github.com/confident-ai/deepeval/releases
  4. DeepEval Test Cases https://deepeval.com/docs/evaluation-test-cases
  5. DeepEval LLM Tracing https://deepeval.com/docs/evaluation-llm-tracing
  6. DeepEval Unit Testing in CI/CD https://deepeval.com/docs/evaluation-unit-testing-in-ci-cd
  7. DeepEval Metrics Introduction https://deepeval.com/docs/metrics-introduction
  8. DeepEval Classifiers https://deepeval.com/docs/classifiers-introduction
  9. DeepEval Synthetic Data Generation https://deepeval.com/docs/synthetic-data-generation-introduction
  10. DeepEval Integrations https://deepeval.com/integrations
  11. Confident AI Pricing https://www.confident-ai.com/pricing
  12. DeepEval Data Privacy https://deepeval.com/docs/data-privacy
  13. AgentCore Third-party Evaluators https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/third-party-evaluators.html
  14. Rubin Lake Technology Radar: DeepEval https://www.rubinlake.com/en/technology-radar/data-platforms-and-mlops/deepeval
  15. PromptLayer: DeepEval Review https://www.promptlayer.com/blog/deepeval-review-what-it-is-and-the-best-alternatives-in-2026/
  16. DeepTeam GitHub https://github.com/confident-ai/deepteam
  17. DeepTeam PyPI https://pypi.org/project/deepteam/
  18. DeepTeam Red Teaming Introduction https://www.trydeepteam.com/docs/red-teaming-introduction
  19. DeepTeam Adversarial Attacks https://www.trydeepteam.com/docs/red-teaming-adversarial-attacks
  20. DeepTeam Risk Assessment https://www.trydeepteam.com/docs/red-teaming-risk-assessment
  21. DeepTeam Code Scanning https://www.trydeepteam.com/docs/code-scanning-introduction
  22. DeepTeam Guardrails https://www.trydeepteam.com/docs/guardrails-introduction
  23. DeepTeam Vulnerabilities https://www.trydeepteam.com/docs/red-teaming-vulnerabilities
  24. DeepTeam Frameworks https://www.trydeepteam.com/docs/frameworks-introduction
  25. DeepTeam Data Privacy https://www.trydeepteam.com/docs/data-privacy
  26. Ragas GitHub https://github.com/vibrantlabsai/ragas
  27. promptfoo GitHub https://github.com/promptfoo/promptfoo
  28. Inspect AI GitHub https://github.com/UKGovernmentBEIS/inspect_ai
  29. LangSmith Pricing https://www.langchain.com/pricing
  30. Langfuse Pricing https://langfuse.com/pricing
  31. Arize Phoenix License https://arize.com/docs/phoenix/self-hosting/license
  32. TruLens GitHub https://github.com/truera/trulens
  33. MLflow GenAI Evaluation https://mlflow.org/docs/latest/genai/eval-monitor/
  34. Opik GitHub https://github.com/comet-ml/opik
  35. Braintrust Pricing https://www.braintrust.dev/pricing
  36. Evidently OSS vs Cloud https://docs.evidentlyai.com/faq/oss_vs_cloud
  37. Agent and model evaluations in Gemini Enterprise Agent Platform are now GA https://developers.googleblog.com/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/
  38. AgentCore Evaluations GA https://aws.amazon.com/about-aws/whats-new/2026/03/agentcore-evaluations-generally-available
  39. Microsoft Foundry Built-in Evaluators https://learn.microsoft.com/en-us/azure/foundry/concepts/built-in-evaluators
  40. OpenAI API Deprecations https://developers.openai.com/api/docs/deprecations
  41. inspect_evals GitHub https://github.com/UKGovernmentBEIS/inspect_evals
  42. Langfuse Cloud Japan https://langfuse.com/blog/2026-04-27-langfuse-cloud-japan
  43. MLflow Third-party Scorers https://mlflow.org/blog/third-party-scorers/
  44. Comet Pricing https://www.comet.com/site/pricing/
  45. Braintrust Run in CI https://braintrust.dev/docs/evaluate/run-in-ci
  46. promptfoo Red Team Plugins https://www.promptfoo.dev/docs/red-team/plugins/
  47. PyRIT GitHub https://github.com/microsoft/PyRIT
  48. garak GitHub https://github.com/NVIDIA/garak
  49. Giskard OSS GitHub https://github.com/Giskard-AI/giskard-oss
  50. PurpleLlama GitHub https://github.com/meta-llama/PurpleLlama
  51. promptfoo Pricing https://www.promptfoo.dev/pricing/
  52. OpenAI to acquire Promptfoo https://openai.com/index/openai-to-acquire-promptfoo/
  53. Moving from OpenAI Evals to promptfoo https://developers.openai.com/cookbook/examples/evaluation/moving-from-openai-evals-to-promptfoo
  54. Langfuse joins ClickHouse https://langfuse.com/blog/announcing-acquisition
  55. Dynatrace to acquire Arize https://www.dynatrace.com/news/press-release/dynatrace-to-acquire-arize/
  56. Splunk / Galileo https://www.splunk.com/en_us/about-splunk/acquisitions/galileo.html
  57. CoreWeave completes acquisition of Weights & Biases https://coreweave.com/blog/coreweave-completes-acquisition-of-weights-biases
  58. Check Point acquires Lakera https://www.checkpoint.com/press-releases/check-point-acquires-lakera/
  59. Palo Alto Networks completes acquisition of Protect AI https://www.paloaltonetworks.com/company/press/2025/palo-alto-networks-completes-acquisition-of-protect-ai
  60. LLM Guard GitHub https://github.com/protectai/llm-guard
  61. Guardrails AI joins Harvey https://www.harvey.ai/blog/guardrails-ai-joins-harvey
  62. NeMo Guardrails GitHub https://github.com/NVIDIA-NeMo/Guardrails
  63. Azure AI Red Teaming Agent https://learn.microsoft.com/azure/ai-foundry/concepts/ai-red-teaming-agent
  64. Nejumi LLM Leaderboard https://nejumi.ai
  65. Ragas Metrics Language Adaptation https://docs.ragas.io/en/stable/howtos/customizations/metrics/metrics_language_adaptation/
  66. Microsoft Foundry Evaluation Regions and Limits https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-regions-limits-virtual-network
  67. LangSmith Regions FAQ https://docs.langchain.com/langsmith/regions-faq
What else is there besides DeepEval? Looking for LLM app evaluation alternatives, I found 2026 was full of acquisitions (promptfoo went to OpenAI, Langfuse to ClickHouse)
https://yurudeep.com/posts/deeplearning/2026/20261005/en/
作者
ひらノルム
公開日
2026-10-05
ライセンス
CC BY-NC-SA 4.0