
opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
The Lens
By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.
Updated Sep 2026
Opik records what your AI app actually did. Every model call, agent step and tool call lands in a trace you can click through, and the same platform runs evals: test datasets, experiments, and AI-graded scores for things like hallucination. It comes from Comet, the ML experiment-tracking company, and the whole repo is Apache 2.0 with no enterprise folder.
Self-hosting is real work. The local install is one script, but it starts about eight containers (ClickHouse, ZooKeeper, MySQL, Redis, MinIO, a Java backend, a Python evaluator and the UI), and Comet says that Docker setup is not for production. Production means Kubernetes and ClickHouse nodes that want 24GB of RAM or more. Anonymous usage stats are on by default; one environment variable turns them off.
The cloud Free plan covers 25,000 spans a month for up to 10 people. Pro is $19 a month for up to 50 people and 100,000 spans, then $5 per extra 100,000. Solo and small teams: start on the free cloud. Teams that can't send prompts to a vendor: self-host. SSO, RBAC and compliance paperwork are Enterprise only. langfuse/langfuse and Arize-ai/phoenix are the closest alternatives, and Phoenix is source-available, not open source.
The catch: self-hosted Opik has no login at all. Comet keeps authentication in its enterprise product, so anyone who can reach the port can read every trace, including your prompts and any user data inside them. Put it behind a VPN or oauth2-proxy/oauth2-proxy first. It also ships almost daily, so plan on frequent upgrades.
Free vs Self-Hosted vs Paid
free self hosted paid cloudFree: Apache 2.0, with no enterprise directories in the repo. Self-hosted gets unlimited spans and retention plus tracing, evals, prompt management, online eval rules, dashboards, alerts, the Agent Optimizer and the PII/topic guardrails (guardrails are self-hosted only). Comet Cloud Free: $0, up to 10 team members, 25k spans/month, 60-day retention.
Self-hosted: Docker Compose for local use (about eight containers: ClickHouse, ZooKeeper, MySQL, Redis, MinIO, Java backend, Python evaluator, frontend), which Comet says is not for production. Production is a Helm chart on Kubernetes with the Altinity ClickHouse operator; Comet's sizing guide puts production ClickHouse at 6 vCPU / 24 GiB and up. No authentication or user management in the open source build: secure it at the network layer. Anonymous usage stats are on by default (OPIK_USAGE_REPORT_ENABLED=false turns them off).
Paid: Pro Cloud is $19/month for up to 50 members (the plan card doesn't list a per-seat price), 100k spans/month, extra spans $5 per 100k, longer retention (60 to 400 days) $29 per 100k spans. Enterprise is custom: managed or on-prem deployment, RBAC, SSO (OAuth, SAML, LDAP), service accounts, SOC 2 / ISO 27001 / HIPAA, custom data region, 2-hour support SLA. OpikAssist AI debugging and the Ollie coding-harness actions are cloud only.
Free to self-host with every feature except login. The $19/month Pro cloud is cheap; SSO and RBAC mean Enterprise.
What to do by team size
- Solo
- free; the cloud Free plan's 25k spans a month is plenty to start
- Small team
- free cloud, or Pro at $19/month once you pass 25k spans
- Medium team
- free self-hosted if you can run ClickHouse and put it behind an auth proxy; otherwise Pro
- Large team
- Enterprise for SSO, RBAC and compliance, or self-host with your own access layer
Get tools like this every Wednesday
One featured tool, three on the radar. No fluff.
Similar Tools

The LLM Evaluation Framework

Open source AI/ML lifecycle platform

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

Open source LLM engineering platform

AI Observability & Evaluation

Easy-to-use & supercharged open-source experiment tracker.
A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores
Trust Signals
License: Apache License 2.0
Use freely. Patent grant included.
Commercial use: ✓ Yes
About
- Owner
- Comet (Organization)
- Stars
- 22,315
- Forks
- 1,834
Explore Further
More tools in the directory
everything-claude-code
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
270.3k ★hermes-agent
The agent that grows with you
250.4k ★ollama
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
182.0k ★