I design, deploy, and hand over production-grade open-source observability stacks — Loki, Mimir, Tempo, Grafana & OpenTelemetry — so your SaaS or AI product gets full logs, metrics, and traces without a per-GB vendor bill.
Datadog, New Relic, and similar platforms charge per host, per GB ingested, per metric series. As you grow, your bill grows faster than your infrastructure. A self-hosted, open-source stack decouples the two.
OpenTelemetry instrumentation feeds a single Collector, which fans out logs, metrics, and traces to purpose-built Grafana Labs storage engines — all queried and correlated from one Grafana instance.
// deployed via Helm on your Kubernetes cluster, or docker-compose for smaller footprints
Every engagement ends with your team owning and operating the stack — documented, dashboarded, and alerting on what matters.
Architecture sized to your traffic and budget: Loki, Mimir, Tempo, Grafana on Kubernetes or docker-compose, backed by S3-compatible object storage.
Structured cutover from Datadog / New Relic / Splunk with dashboard and alert parity, run side-by-side until you're confident to switch.
Auto- and manual-instrumentation across your services in Go, Python, Node, Java — consistent trace context propagation end-to-end.
SLO-driven Grafana dashboards and alert rules routed to Slack/PagerDuty — built around what actually pages someone, not vanity charts.
Find where your current bill (or your new stack's resource usage) is going — label cardinality, retention policy, and sampling review.
Architecture docs, upgrade playbooks, and a live walkthrough with your engineers so the stack is fully yours to operate.
If you're building on top of LLMs, "it feels slow" and "the bill went up" aren't answerable questions without instrumentation. I wire up OpenTelemetry's GenAI semantic conventions so every model call becomes a structured, queryable span.
Scope and timeline flex with the size of your environment — this is the shape of a typical engagement.
Review current tooling, spend, traffic volume, and compliance constraints. Define target SLOs and what "good" observability looks like for your team.
Loki, Mimir, Tempo, and Grafana deployed to your infrastructure with object storage, retention, and multi-tenancy configured to spec.
Instrument services and (if applicable) LLM/AI call paths. Wire logs, metrics, and traces together with consistent resource attributes.
Build the dashboards and alert rules your team will actually use daily, plus documented runbooks for common incidents.
Live walkthrough with your engineers, recorded for future hires, plus a 30-day post-handover support window.
Send a short note about your current stack, traffic, and what's driving the change — I'll reply with next steps within one business day.