Use case · Platform & infra

One observability story across every team shipping agents.

You support many product teams, many runtimes, many model vendors. Trefur gives you one canonical trace format — built on OpenTelemetry — that every team can adopt without changing framework, vendor, or backend.

What you are trying to do.

  • Give every product team a paved road for shipping agents, without dictating which framework they pick.
  • Stand up one telemetry pipeline that survives the next framework change of the year.
  • Roll out a shared alerting story — Slack, PagerDuty, webhook — without each team rebuilding it.
  • Keep cost visibility per team, per agent, per environment without chasing receipts.
  • Stay neutral on model vendors. Teams will swap OpenAI for Anthropic for Bedrock and back.

What hurts today.

Three teams, three telemetry pipelines

Team A wired LangChain callbacks to Datadog. Team B is on a custom logger to S3. Team C posts JSON blobs to a Slack channel. Asking a cross-team question takes a week.

OTel is great until the agent layer

Your service spans are tidy. Your agent spans are a wall of free-form log lines because nobody emits structured tool calls. You cannot graph what you cannot name.

Vendor lock-in by accident

Whichever observability product a team picked first wired itself into the framework. Swapping it now means losing six months of dashboards.

Cost reporting is a manual roll-up

Finance asks who is driving the OpenAI bill. You spend a Friday joining model invoices to per-team logs by hand. The numbers are stale before the meeting.

What Trefur gives you.

One canonical agent trace

OpenTelemetry GenAI semantic conventions on the wire. Every team that emits to the standard gets a real trace, regardless of framework.

Fan-out friendly

Trefur is a second OTel exporter, not a replacement. Teams keep Datadog, Honeycomb, Grafana — and finally have one agent-aware view across them.

Self-hostable collector

Single Go binary in your network. Local batching, on-disk buffering, redaction. Routes to Trefur and your existing backend in parallel.

Per-team cost and latency

Roll up traces by team, agent, environment, model vendor. Answer the finance question with a chart, not a spreadsheet.

Alert routing by ownership

Route alerts to Slack, PagerDuty, Teams, or webhook per team or per agent. No more shared inboxes that nobody owns.

Open standards keep you neutral

Adopt LangChain, CrewAI, AutoGen, Pydantic AI, MCP, or your own orchestrator. Trefur reads them all from the same OTel surface.

A week in the life with Trefur.

  1. Mon

    You publish a short paved-road doc: install the collector, point your SDK at it, traces show up in Trefur. Two product teams are migrated by end of day.

  2. Tue

    Finance asks for last week's model spend by team. You send them a Trefur link filtered by team tag, broken down by model vendor. Five minutes.

  3. Wed

    A third team is on a homegrown orchestrator. They emit OTel GenAI spans directly. No SDK install, no framework switch — traces appear.

  4. Thu

    You set a baseline alert: any agent whose retry rate doubles week-over-week pages the owning team in PagerDuty with the worst trace already linked.

  5. Fri

    Quarterly review. You show one dashboard: agents shipped, runs per day, cost per run trending down 18% as teams tuned their prompts using the trace data.

Make agent observability a paved road.

Free tier for the first traces. Team and Enterprise scale with volume, retention, and support.