# honeyhive.ai > AI-optimized mirror of honeyhive.ai containing 33 pages totalling 6,241 words of clean markdown content, structured data, and semantic HTML. Original source: https://honeyhive.ai. Last updated: 2026-07-20T14:37:42.012Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [HoneyHive - The Observability Layer for Production Agents](/content/site-root.html): HoneyHive unifies observability and evaluation into a continuous improvement loop, so every team across your business can ship quality agents with confidence. (833 words) ## Articles & Blog Posts - [Privacy Policy](/content/privacy/index.html): HoneyHive Inc. Privacy Policy Agreement (3,128 words) - [HoneyHive - Evaluation](/content/evaluation/index.html): HoneyHive's AI evaluations platform helps developers automatically test, evaluate, and simulate AI agents against datasets using code metrics, LLM-as-a-judge, or human review. (532 words) - [The Tokenomics Problem With Coding Agents](/content/post/the-tokenomics-problem-with-coding-agents/index.html): Coding-agent costs are exploding. Trace spend by session, step, model, and outcome, so you can fix the runs that waste tokens instead of only capping them. (41 words) - [HoneyHive - Prompt Management](/content/playground/index.html): HoneyHive prompt management platform enables your entire team to manage, version, and deploy new prompts and models within a unified, collaborative workspace. (276 words) - [Pricing](/content/pricing/index.html): Free for individual developers. Powerful capabilities for scaling teams. Dedicated support and self-hosting for large enterprises. (365 words) - [Research and Insights](/content/blog/index.html): Stay up to date with the latest research, insights and announcements from HoneyHive. (141 words) - [Introducing HoneyHive v2](/content/post/introducing-honeyhive-v2/index.html): v2 is a ground-up refactor of the platform, featuring a new architecture, Custom Roles, new Python and TypeScript SDKs, HoneyHive CLI, Trajectories for long-running agents, and more. (39 words) - [Trace and evaluate Microsoft Copilot Studio agents with HoneyHive](/content/post/trace-and-evaluate-microsoft-copilot-studio-agents-with-honeyhive.html): HoneyHive now brings full tracing and evaluation to Microsoft Copilot Studio. (28 words) - [Open-Source vs OpenAI: Is it Time to Move On?](/content/post/openai-vs-open-source-models/index.html): After recent OpenAI reliability issues and uncertainty surrounding the company's future, many enterprises are looking to move towards open-source AI but find themselves unprepared for the switch. The solution is simple - start your data flywheel early. (54 words) - [Our Open-Source Model Selection Guide](/content/post/model-selection-guide/index.html): This guide serves as an overview of the open-source LLM ecosystem and breaks down how you should select the right open-source model based on your use-case and enterprise-specific requirements. (44 words) - [What LLM Benchmarks Can and Cannot Tell You](/content/post/what-llm-benchmarks-can-and-cannot-tell-you/index.html): LLM benchmarks abound, but the devil's in the details. We'll unpack their strengths, pitfalls, and how to read between the lines. (40 words) - [Product Update: Offline Evaluations](/content/post/product-update-offline-evaluations/index.html): We're introducing powerful new updates to our Offline Evaluations product, enabling developers to thoroughly test and improve their GenAI applications. (33 words) - [The Agent Development Lifecycle](/content/post/the-agent-development-lifecycle/index.html): The playbook for shipping reliable agents across teams—again and again. (23 words) - [The Evolution of Observability: From Monoliths to AI Agents](/content/post/the-evolution-of-observability-from-monoliths-to-ai-agents.html): Software and ML Observability struggle to scale to the world of AI agents. How did we get here? What’s the way forward? (39 words) - [Towards Evaluation Driven Development with MongoDB and HoneyHive](/content/post/towards-evaluation-driven-development-with-mongodb-and-honeyhive.html): The Step-by-Step Guide to Build Production-Ready RAG Systems with MongoDB and HoneyHive (29 words) - [Introducing Annotation Queues](/content/post/introducing-annotation-queues/index.html): We're introducing Annotation Queues - a new way to scale human judgement and domain expertise via in-app human review. (31 words) - [Scale Agent Governance with Microsoft's ASSERT and HoneyHive](/content/post/scale-agent-governance-with-microsofts-assert-and-honeyhive.html): Microsoft's ASSERT turns your policies into generated evals that scale governance. HoneyHive traces every session so your team can review and annotate evals alongside your production traces. (43 words) - [Product Update: Traces, Datasets, and Online Evaluation](/content/post/introducing-rag-agent-evaluation/index.html): Today, we are launching Datasets, unveiling a brand-new data model optimized for distributed tracing, releasing advanced capabilities that enable agent and RAG evaluators, introducing new monitoring chart types, and announcing VPC deployment options., Click through a step-by-step, interactive demo walkthrough of Honeyhive, powered by Supademo., Click through a step-by-step, interactive demo walkthrough of Honeyhive, powered by Supademo., Click through a step-by-step, interactive demo walkthrough of Honeyhive, powered by Supademo. (49 words) - [How to Escape the Eval Cold Start Problem](/content/post/how-to-escape-the-eval-cold-start-problem/index.html): Get your eval flywheel spinning as fast as possible. (27 words) - [How to Evaluate Compound AI Systems](/content/post/how-to-evaluate-compound-ai-systems/index.html): Learn how to evaluate complex multi-step AI systems by breaking them down into measurable components. Using a RAG pipeline example, discover how tracing and component-level metrics can help build more reliable AI products. (48 words) - [Tracing RAG applications in production with LanceDB and HoneyHive](/content/post/moving-ai-applications-to-prod-with-lancedb-and-honeyhive.html): Learn how to trace and evaluate AI applications at scale with HoneyHive and LanceDB. (32 words) - [Introducing Role-Based Access Control](/content/post/introducing-role-based-access-control/index.html): We're introducing Role-Based Access Control with two-tier permissions to secure enterprise AI applications, prevent data leakage between teams, and ensure regulatory compliance. (35 words) - [How to evaluate AI applications](/content/post/evaluating-ai-apps/index.html): This guide introduces key steps, metrics, and techniques you need know to set up reliable, industrial-grade testing for AI applications. (34 words) - [Introducing HoneyHive Skills](/content/post/honeyhive-skills/index.html): Instrument, Evaluate, and Improve Your AI Agents Faster. (20 words) - [HoneyHive Recognized in the 2026 Gartner® Market Guide for AI Evaluation and Observability Platforms](/content/post/honeyhive-recognized-in-the-2026-gartner-market-guide-for-ai-evaluation-and-observability-platforms.html): HoneyHive is recognized by Gartner as an emerging leader on their 2026 Market Guide for AI Evaluation and Observability Platforms. (42 words) - [Search That Learns From You: Building Adaptive Retrieval Systems using Qdrant & HoneyHive](/content/post/building-context-aware-systems-with-qdrant-honeyhive.html): This guide demonstrates an interactive approach that leverages positive and negative feedback to continuously refine search results using Qdrant's Discovery API and HoneyHive's observability platform. (47 words) - [Evaluations Across the Agent Development Lifecycle](/content/post/evaluations-across-the-adlc/index.html): A stage-by-stage guide to applying the right evaluation rigor across the agent development lifecycle. (28 words) - [HoneyHive achieves SOC 2 Type II, GDPR, and HIPAA compliance](/content/post/honeyhive-achieves-soc-2-type-ii-gdpr-and-hipaa-compliance.html): We're excited to announce HoneyHive has achieved SOC 2 Type II, GDPR, and HIPAA compliance! (34 words) - [Standardizing AI Observability Before It Breaks: A Case Study on 73,000 Agent Schemas](/content/post/agent-schema-case-study/index.html): Why AI observability needs a standard — and why we're betting on OTel GenAI. (37 words) - [Avoiding Common Pitfalls in LLM Evaluation](/content/post/avoiding-common-pitfalls-in-llm-evaluation/index.html): Discover the hidden challenges of LLM evaluation and the most common mistakes we've seen after helping hundreds of teams build effective evals that drive business results. (41 words) - [Announcing HoneyHive](/content/post/announcing-honeyhive/index.html): Introducing HoneyHive - The tool to iterate and optimize your LLM-powered products (23 words) - [2025: A Year in Review](/content/post/2025-year-in-review/index.html): What we shipped in 2025 to make production AI agents actually work. (25 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/robots.txt): Crawler directives