
Agenta
Open-source LLMOps platform for prompts, evals, and observability
4.6k stars 641 forks last commit first released
Actively maintained
Last commit 26 Aug 2026.

Agenta is an open-source LLMOps platform for building and operating production-grade LLM applications. It centralizes prompt work, evaluation workflows, and runtime traces so teams can iterate safely and measure quality over time.
Key Features
- Interactive prompt playground to compare prompts and models side-by-side on real test cases
- Prompt and configuration versioning with environments/branching to control changes
- Testset management (including CSV import and capturing production cases) for repeatable experiments
- Automated and human evaluation workflows, including LLM-as-judge and custom evaluators
- Production observability with tracing, latency/usage/cost tracking, and debugging via detailed traces
- Open standards support for tracing via OpenTelemetry-compatible instrumentation
- UI and API parity to support both expert-driven and engineering workflows
Use Cases
- Prompt engineering and regression testing before shipping changes to production
- Evaluating agents and RAG pipelines with automated metrics plus expert review
- Debugging and monitoring production LLM apps to detect failures and performance regressions
Agenta fits teams that need a single source of truth for prompts, evaluations, and traces, combining experimentation and operational monitoring in one platform. It helps reduce trial-and-error iterations by making changes measurable and auditable.
Categories:
Tags:
Tech Stack:
Similar to Agenta

Langfuse
Open-source platform for LLM observability, evals, and prompt management
Langfuse is an open-source LLM engineering platform for tracing, metrics, evaluations, datasets, and prompt management to debug and improve AI applications.

Opik
LLM observability and evaluation platform for traces, tests, and dashboards
Opik is an open-source platform to trace, evaluate, and monitor LLM apps, RAG pipelines, and agent workflows with automated evaluations and production dashboards.
Scriberr
Offline AI audio and video transcription with transcript chat
Scriberr is a self-hosted, privacy-focused AI transcription app for audio and video, with speaker diarization, word-level timestamps, summaries, and transcript chat.
Perplexica
Privacy-focused AI answering engine with web search and citations
Self-hosted AI answering engine that combines web search with local or hosted LLMs to generate cited answers, with search history and file uploads.

Ollama
Run and manage large language models locally with an API
Ollama is a local LLM runtime that lets you pull, run, and customize models, offering a CLI and REST API for chat, generation, and embeddings.

AnythingLLM
All-in-one AI chat app with RAG, agents, and multi-model support
AnythingLLM is an all-in-one desktop and Docker app for chatting with documents using RAG, running AI agents, and connecting to local or hosted LLMs and vector databases.



