epok
Recognize. Recover. Resolve.

See it first.
Recover with a fix you can reverse.
Make sure it never returns.

One engine across all three tenses of an incident — it catches the failure from your first line of logs, with no rules to write, recommends a reversible fix to restore service, then drafts the cited root cause so it doesn't recur — across logs, metrics, traces, infrastructure, and RUM.

Try the real product Every claim cited to the line When the evidence is thin, it says so Your telemetry never trains a model
The Epok incident workspace: a checkout saturation cascade with the Before/During/After arc, a reversible recovery recommendation, and a cited root-cause verdict at 95% calibrated confidence the live demo — this exact screen, no signup →
Before → During → After

One incident, three tenses — one engine.

One real incident, worked end to end — a retry storm holding a service down 14 minutes after its trigger cleared. The same incident is open in the live demo.

IMMINENTqueue_utilization · checkout
now 81%saturation in ~12m
BEFORE · Recognize
Recognize

The request queue is accelerating toward saturation. Epok forecasts the breach, with live ETA and confidence shown, 12 minutes before anything pages.

RECOVER· we recommend — you execute
Shed load: disable client retries
UNDO · EASYBLAST · MEDIUMCONFIDENCE · HIGH
⚠ Rollback won't recover this — nothing failing was deployed.
DURING · Recover
Recover

The counterintuitive right move: shed the retries. And the guard that matters — rollback will NOT recover this. We recommend; you execute.

VERDICT· cited · filed ✓
TRIGGER · CLEARED40s packet loss, self-resolved
SUSTAININGretry policy — ×4.2 offered load
last deploy exonerated · every claim linked to its line
AFTER · Resolve
Resolve

The verdict separates the trigger (a 40s blip, cleared) from the sustaining cause (the retry policy) — cited, with the last deploy exonerated.

before · during · after — one incident, worked end to end · 2:00
THE VERDICT THAT ADMITS WHEN IT DOESN'T KNOW

It commits when the evidence is there. When it isn't, it says so.

Every other tool answers with confidence whether or not it has any. Epok abstains — and tells you exactly why. An honest "not enough evidence" is worth more than a fluent guess you have to disprove at 3am. No incumbent ships a verdict that knows the difference.

COMMITTED
Root cause: checkout-service CPU saturation, sustained 97%. Cited to the metric + the two spans that corroborate it.
ABSTAINED
"Three candidates are close and none dominates — committing to one would be a guess, and we don't guess." Here are the three, ranked, with what each is missing.
The demo publishes its own scorecard — committed vs abstained, last 90 days, on every verdict. See the live track record →
the shippers you already run, pointed at one endpoint — no agent of ours to install, no SDK to embed
0
rules to write

Detection runs automatically across your signals. On every tier, including the trial.

<5min
to first alert

Point any shipper at one URL. Detection starts on the first matching line.

100%
claims cited

Each AI root cause links the exact log line that produced it. No narrative without proof.

The wedge · one engine, three tenses

Everyone else starts the clock when it breaks. Epok works before, during, and after.

SEARCH-FIRST

You go looking for the problem.

Fast search and an assistant that summarizes results — but only when you ask. You have to know what to query; nothing fires until someone opens it.

$ search "error" | filter svc=payment
EPOK · THE WHOLE ARC

The problem finds you — early.

Before it pages, Epok forecasts the breach with lead time. During, it recommends the reversible fix to restore service. After, it drafts the root cause — every claim cited.

Resolve · one investigation surface

The whole incident on one screen. No tab-hopping.

When an alert fires, Epok opens one canvas: the drafted root cause, what changed, the cascade timeline, the blast radius, and who's affected — every claim clickable back to the line, span, or metric behind it.

An Epok investigation on one canvas: a cited probable-cause verdict with a confidence score, blast radius across services and users, and the cross-service cascade timeline — every claim links to its evidence
read-only · every claim clickable back to the line, span, or metric · open it on the live demo →
session_replay · stitched to the trace

Watch the user hit the bug — from the trace that broke

Every session replay is linked to the backend trace by a shared trace ID. From a failing trace on this canvas, one click opens the recording of the exact user who hit it — seeked to the moment of failure, the lead-up a scrub away. Not two tools you cross-reference by timestamp. One click.

trace 69e9…2be4 · POST /checkout · 500▶ replay · seeked to 00:04 — "Payment failed"
RUM & Replay →
customer_impact

Customer impact, on every alert

Paste a roster (id, name, tier). Every alert scans the incident window for customer_id and joins to the roster. Your VP of CS reads the same alert as your on-call.

0Enterprise0Pro0Free
wire it — paste a roster, done →
ask_epok

Plain English → search

Type "why is checkout slow in the last hour." AI translates it to a query, a 42-test validator forces time + limit, then we run it. Query is ground truth; the explanation sits beside it.

› why is checkout slow in the last hour
→ search svc=checkout p95>1s | last 1h | limit 500
postmortem

Postmortem draft, the moment it resolves

When the incident closes, a draft appears — triggering signal, cited evidence chain, matched playbook, and customer-impact rollup, pre-assembled. You edit; you don't author from a blank page.

triggerevidenceplaybookimpact
playbook_match

Your runbooks, matched — not authored

Bulk-import from Confluence, Notion, GitHub, or Markdown. A citation engine surfaces the best-matching runbook — the specific steps land inside Slack, PagerDuty, and the deep RCA, not a link to a wiki.

payment-pool exhaustion → 3 steps inlined in PagerDuty
Recognize · what it catches

Catch what actually pages you.

An error you've never seen

Surfaces messages that never appeared in your recent history — the first sign of a fresh failure.

payment-service: "FATAL: connection pool exhausted" — first seen
COLD-START · ready within your first week

A service gone quiet

Catches a service that stops logging when it normally logs steadily. Failure with no error — just absence.

worker-billing went silent — last log 6m ago (normally 30s)
COLD-START · ready within your first week

Spikes, drops, flatlines

Flags volume that jumps, falls, or flatlines against each service's daily and weekly normal.

api: 12,400 lines/min vs 3,200 normal (× 3.9)
COLD-START · learns your normal

Many errors, one root

Groups errors that share a shape, so dozens of variants land as a single alert.

84 variants of one failure folded into 1 alert
COLD-START · active immediately

Crashing workloads

Surfaces crash loops, out-of-memory kills, and unschedulable workloads straight from their logs.

billing-7c4b out-of-memory — 3rd restart in 4m
COLD-START · fires from minute one

Failures that cascade

Connects upstream failures, retries, and circuit-breaker trips into the cascade they cause.

3 services blame one upstream — cascade in 8s
COLD-START · fires from minute one

One incident. Not fifty alerts.

The pager is rationed by design. Repeats collapse, cascades arrive as one chain, and severity rides explicit thresholds the product enforces.

01

Fingerprint dedup

The same root cause never pages twice — repeats collapse into one alert with a fire count.

02

Incident grouping

A cascade arrives as one page with the full chain: db silent → API refused → frontend 502s.

03

Dynamic suppression

Repeat fires de-escalate on a widening window. New shapes still get full severity.

04

Severity rationing

Critical / Warning / Info on explicit thresholds the product enforces.

Live demo · no signup

See it on data. No signup.

A 5-service app generates a continuous synthetic log stream into a public Epok tenant. Anomaly detection, RCA, and clustering run on it live — Epok working on real-shape data, not a marketing video.

alerts inboxlive
CRIT
connection pool exhausted
payment-service · first seen · 3-svc cascade
WARN
service went silent — 6m
worker-billing · normally logs every 30s
CRIT
out-of-memory — 3rd restart
billing-7c4b · crash loop · 4m
INFO
12,400 lines/min vs 3,200
api-gateway · × 3.9 normal
Pricing

Flat price. No surprise bill. No cardinality tax.

One meter for every signal — logs, metrics, traces, infrastructure, RUM, and replay on one bill. No per-host, per-query, or cardinality charges.

Trial
$0 / 14 days
  • Up to 1 TB · full retention
  • Every feature unlocked
  • No credit card
Start free →
Growth
$599 / mo flat
  • 4 TB / month · 30-day retention
  • Unlimited users · SSO
  • Priority support
Start Growth →
FAQ

Before you ask.

Ingest pauses, your data stays readable, and you add a card when you're ready. Nothing auto-charges.

No. A flat monthly price covers your included volume; overage, if you ever exceed it, is $0.20/GB, posted to your dashboard daily — never a surprise.

No. Log as many unique fields as you want; there is no per-series tax.

No. Detectors run automatically, and you can ask in plain English.

Logs, metrics, traces, infrastructure, RUM, and session replay — point any shipper at one URL. Many ingest formats — no agent of ours to install, no SDK to embed. Run it alongside your current tool during the trial.

NO CREDIT CARD · 30 SECONDS

Point your stack at Epok.
Get the cause back.

$ curl -X POST https://in.epok.dev/v1/logs \
-H "Authorization: Bearer $EPOK_KEY" \
-d '{"service":"api","level":"info","msg":"hello"}'
# {"ok":true}