Skip to main content
Production intelligence / evidence-first incident response Self-hosted · approval-gated

Every conclusion carries its evidence

Production Master investigates your production incidents and returns a report where every claim cites the log, metric or diff it came from — and nothing is changed without your approval.

Self-host via Helm — available now to closed-beta teams.

Evidence workspace · illustrative data
PM-SAMPLE-001checkout-api — p99 latency regressionAwaiting human approval

Conclusion

The connection-pool ceiling was cut from 80 to 20 in deploy a83f9c, starving checkout-api under normal traffic.

p99 410 ms 2.9 s · onset 14:23 UTC · seconds after deploy

Evidence

E-01
p99 request latency stepped from 410 ms to 2.9 s and held.
metrics · http_request_duration_seconds · 14:23:11 UTC
E-02
DB_POOL_MAX reduced 80 → 20 in the deployed values file.
deploy diff · checkout-api@a83f9c · 14:23:14 UTC
E-03
Pool exhaustion in application logs: 0 idle connections at the new limit, timeout_ms=2000.
logs · checkout-api · 2,140 matching lines · 14:23:19 UTC
E-04
Cache hit rate within baseline and upstream latency unchanged across the window.
metrics · cache and upstream telemetry · 14:27:42 UTC
Considered · rejectedCache degradation

Cache degradation does not match the evidence — the hit rate held within baseline through the entire regression window [E-04], and no cache-layer change was deployed. Discarded before the conclusion was formed.

Confidence

87%

Sample confidence. The conclusion rests on 3 cited exhibits, each naming its own source.

Proposed action

Restore the previously reviewed DB_POOL_MAX value, then observe p99 and pool errors for ten minutes.

Human approval required

Production Master will not apply this change itself. It waits here.

Sources consulted

  • deploy history1 commit
  • metrics3 series
  • application logs2,140 lines
  • cluster events18 events

See the product, not a promise

One run, start to gate

A minute and a half of the real interface, on a loop. Scrub it, pause it, or open it full screen — evidence gathered, the approval gate, and the fix, in the order they happen.

Read how the run worksReplay a recorded public issue
Reads the evidence you already produce
deploy diffsmetric seriesapplication logspod & node eventsconfig changestraces

An answer you can audit

Three questions decide whether an incident report is worth reading. Production Master answers them with the conclusion, the evidence under it, and the source behind every exhibit.

01

“How would it know?”

Gathers, then cites

Every log line, metric window, deploy diff and config change it reads becomes a cited exhibit. Nothing enters the report uncited.

claim → E-01E-03 → raw source

02

“What if it’s confidently wrong?”

Shows what it rests on

The explanation it settles on is recorded with the exhibits that support it, so you can check the reasoning instead of trusting the verdict. PM-SAMPLE-001 keeps its rejected alternatives the same way.

considered · rejected · why

03

“Is it going to touch my production?”

Stops at the gate

It proposes the change and waits. Remediation is a decision a human makes, holding a report they can check line by line.

A-01 proposed → awaiting approval

Run it yourself, or let us run it

The investigation engine is the same either way — what changes is who operates it.

Your infrastructure · your Kubernetes clusterself-host via HelmBeta
Evidence sources
metric series
application logs
deploy diffs & config
pod & node events
Helm release
Investigation runtime
The engine is identical to the managed service. No specific cloud or execution platform is required.
Cited report + proposed action
stays inside the boundary until a person acts
leaves the cluster
LLM provider — your own keys
or ours, on the managed plans
Managed service

We run the same runtime. The evidence contract, the citations and the approval gate do not change with the hosting choice.

On-premisesBeta

Workloads pinned to a region you choose, models imported offline, nightly restore drill against a published budget.

Air-gappedPlanned

A fully disconnected install — local models, packaged upgrades, support commitments — is roadmap work, not shipped.

4 evidence sources, one runtime · shipped, closed beta and planned never share a box

$0 self-hosted with your own LLM keys, $199/mo managed, From $60,000/yr for Enterprise. See what each plan includes

Ship the fix because you checked, not because you guessed

Bring one real incident.
The first cited report lands in week 1 of the 8-week pilot — evidence you can check, not a promise.

Self-host via Helm — available now to closed-beta teams.