Independent Observability Consultant

Own your telemetry.
Stop renting it by the gigabyte.

I design, deploy, and hand over production-grade open-source observability stacks — Loki, Mimir, Tempo, Grafana & OpenTelemetry — so your SaaS or AI product gets full logs, metrics, and traces without a per-GB vendor bill.

60–80%Lower telemetry spend
100%Data stays in your VPC
OTelNative GenAI semconv
trace · checkout-service
api-gateway 42ms auth-svc 24ms inventory-svc 11ms postgres query redis cache llm.chat 40ms logs ⇥ correlated p99 latency (mimir) — 460ms window
The Problem

Vendor observability bills scale with your success — not your value.

Datadog, New Relic, and similar platforms charge per host, per GB ingested, per metric series. As you grow, your bill grows faster than your infrastructure. A self-hosted, open-source stack decouples the two.

Per-GB vendor pricing

  • Costs scale linearly (or worse) with log/metric volume
  • High-cardinality metrics get throttled or billed as overage
  • Data lives in someone else's cloud, exportable on their terms
  • Teams quietly under-instrument to control cost

Self-hosted OSS stack

  • You pay for compute and object storage — not per event
  • Loki/Mimir/Tempo scale horizontally on cheap object storage
  • Full data residency and retention control, inside your VPC
  • Instrument everything — logs, metrics, traces, LLM calls
Architecture

One pipeline, three signals, correlated by design.

OpenTelemetry instrumentation feeds a single Collector, which fans out logs, metrics, and traces to purpose-built Grafana Labs storage engines — all queried and correlated from one Grafana instance.

Your Services LLM Calls Infra / Hosts OpenTelemetry Collector Loki logs Mimir metrics Tempo traces Grafana dashboards + alerts unified correlation object storage (S3 / GCS / MinIO) — cheap, durable, horizontally scalable

// deployed via Helm on your Kubernetes cluster, or docker-compose for smaller footprints

Services

Fixed-scope engagements, not open-ended retainers.

Every engagement ends with your team owning and operating the stack — documented, dashboarded, and alerting on what matters.

Stack Design & Deployment

Architecture sized to your traffic and budget: Loki, Mimir, Tempo, Grafana on Kubernetes or docker-compose, backed by S3-compatible object storage.

Migration Off Vendor Tools

Structured cutover from Datadog / New Relic / Splunk with dashboard and alert parity, run side-by-side until you're confident to switch.

OpenTelemetry Instrumentation

Auto- and manual-instrumentation across your services in Go, Python, Node, Java — consistent trace context propagation end-to-end.

Dashboards & Alerting

SLO-driven Grafana dashboards and alert rules routed to Slack/PagerDuty — built around what actually pages someone, not vanity charts.

Cost & Cardinality Audits

Find where your current bill (or your new stack's resource usage) is going — label cardinality, retention policy, and sampling review.

Runbooks & Handover

Architecture docs, upgrade playbooks, and a live walkthrough with your engineers so the stack is fully yours to operate.

AI / LLM Applications

See inside every prompt — tokens, latency, and cost, per call.

If you're building on top of LLMs, "it feels slow" and "the bill went up" aren't answerable questions without instrumentation. I wire up OpenTelemetry's GenAI semantic conventions so every model call becomes a structured, queryable span.

  • Per-request cost attribution  — trace prompt + completion tokens back to model, route, and customer.
  • Latency breakdown  — separate model TTFT/TTLT from your retrieval, tool calls, and application logic.
  • Prompt & completion capture  — configurable redaction for sensitive fields, correlated with the parent trace.
  • Multi-provider support  — OpenAI, Anthropic, Bedrock, and self-hosted models under one semantic schema.
span.namegen_ai.usagecost
chat.completions
$0.041
embeddings
$0.006
vector.search
$0.000
tool.call:crm
$0.000
chat.completions
$0.058
gen_ai.systemanthropic · openaiΣ $0.105
Engagement Process

From audit to handover, typically 3–6 weeks.

Scope and timeline flex with the size of your environment — this is the shape of a typical engagement.

Week 0 — Discovery

Audit & scoping call

Review current tooling, spend, traffic volume, and compliance constraints. Define target SLOs and what "good" observability looks like for your team.

Week 1–2 — Build

Stack deployment

Loki, Mimir, Tempo, and Grafana deployed to your infrastructure with object storage, retention, and multi-tenancy configured to spec.

Week 2–4 — Instrument

OpenTelemetry rollout

Instrument services and (if applicable) LLM/AI call paths. Wire logs, metrics, and traces together with consistent resource attributes.

Week 4–5 — Operationalize

Dashboards, alerts & runbooks

Build the dashboards and alert rules your team will actually use daily, plus documented runbooks for common incidents.

Week 5–6 — Handover

Knowledge transfer

Live walkthrough with your engineers, recorded for future hires, plus a 30-day post-handover support window.

Let's Talk

Ready to see what your telemetry actually costs?

Send a short note about your current stack, traffic, and what's driving the change — I'll reply with next steps within one business day.