โ€” Dev Tool

Honeycomb

Last updated June 20, 2026 ยท Reviewed by ToolForge Editorial

The observability platform engineers actually like. Now with first-class LLM tracing.

โ˜… 4.7/5 ยท Used by Slack, HelloFresh, LaunchDarkly ยท Since 2016 ยท Free tier (20M events/mo)

Observability for humans, now with LLMs

Honeycomb has been the engineer's choice for production observability since 2016 โ€” Slack, HelloFresh, and LaunchDarkly run on it. In 2025 they shipped native LLM tracing that uses the same OpenTelemetry standard as everything else. Paste in your OpenAI/Anthropic calls, get full traces of every prompt, every tool call, every token cost, every latency spike. And because it's Honeycomb, you get the famous query builder that lets you slice by 50 dimensions in one query โ€” no PromQL, no SQL.

Who it's for: Engineering teams running LLM apps in production who already think in traces, not logs.

Key features

OTel-native OpenTelemetry LLM tracing

Auto-instrumented spans for OpenAI, Anthropic, Cohere, Bedrock. Drop-in SDK, traces show up immediately.

BubbleUp AI-assisted debugging

Click any spike in latency or cost. Honeycomb auto-finds the field that correlates. No manual query writing.

Token cost Per-call spend tracking

Every LLM call shows prompt tokens, completion tokens, and dollar cost. Group by user, model, feature.

Query builder 50-dimensional slicing

No PromQL, no SQL. Drag fields, filter, groupby, visualize. The tool observability should have been all along.

Service map Live dependency graph

See which downstream services, vector DBs, and models your LLM depends on. Spans light up red when one is slow.

Boards Dashboards that ship

Build dashboards in minutes. Share links, embed in Slack. No more "check Grafana" tickets.

The honest take

โœ“ What works

  • OpenTelemetry-native โ€” no vendor lock-in, spans work in any OTel backend
  • BubbleUp AI finds the correlated field in seconds โ€” this is what other tools sell as "AI debugging"
  • Token cost tracking per user/feature is something most teams build themselves
  • Free tier is 20M events/mo โ€” generous enough for small LLM apps
  • Service map shows the full call graph including RAG retrievals, not just model calls

โœ— What doesn't

  • $130/mo Pro is steep if you only need LLM tracing โ€” LangSmith is cheaper for that specific use case
  • Onboarding the OTel SDK takes an afternoon โ€” not as turnkey as LangSmith
  • Free tier requires a credit card after the first 14 days
  • Eval features are basic โ€” for deep eval work, pair with Braintrust or Patronus

Verdict

If you're already an OpenTelemetry shop or you need to observe LLM apps alongside the rest of your microservices, Honeycomb is the obvious pick. The query builder is unmatched, BubbleUp genuinely saves hours, and the LLM tracing is first-class without being a separate product. For pure LLM eval workflows, Braintrust or LangSmith are more focused โ€” but for full-stack observability, nothing else comes close.

๐Ÿ’ก Transparency: This review contains affiliate links. If you sign up through our link, we may earn a commission at no cost to you. We only recommend tools we use ourselves. Full disclosure.

Related Tools

Try Honeycomb today

$130/mo Pro plan

Get Honeycomb โ†’