Zenmem
Login
๐Ÿšจ

On-call incident agent

EngineeringSREPython ยท FastAPI

About this agent

Resolved incidents go in with more than the fix: what was tried and did not work is recorded too, because that is the hour an engineer wastes rediscovering it at 2am. When a new incident is opened, it is matched on symptoms rather than on service name โ€” two different failures on the same service must not be conflated, and a matcher that fires on the service alone is worse than useless when someone is paged. If nothing resembles it, it says so rather than forcing a weak match. Notes are added to the timeline as things happen, so the write-up afterwards is assembled rather than reconstructed from memory.

RUNTIMEPython ยท FastAPI
MEMORY TYPELong-term, per service
SDKzenmem 0.4.4
LOCAL PORT8900

What changed with Zenmem?

The same agent, built twice against the same contract โ€” once on Zenmem, once on MongoDB + LangChain/LangGraph.

Before โ†’ after

Code for the incident record APIโˆ’24%

Code for symptom matchingโˆ’52%

New infrastructure to stand upnone

New dependencies to install0

Schema, collection and index worknone

Matching a new incident on symptomsincluded

Holding history and the service graph togetherone scope

Before With Zenmem

What the team gained

  • Symptom matching is semantic recall over the incident history, so two different failures on the same service stay apart โ€” a name match would have needed a hand-tuned index to beat.
  • The dependency graph and the incident history live in one service scope, so what breaks if this changes is answerable without joining two stores at 2am.
  • What was tried and did not work is stored as freely as what worked, because there is no schema deciding which fields an incident is allowed to have.
  • Each resolution written back improves the next match with no reindexing step to run.
  • An incident with no precedent comes back empty rather than forced into a weak match, since nothing is padding a result set to a fixed size.

How memory is scoped

Long-term memory scoped to a service, holding both the incident history and the service dependency graph โ€” what breaks if this changes. Each resolution is written back, so the third occurrence of a problem is matched faster than the second was. In a live deployment the dependency graph comes from the repository companion rather than being typed in.

How it works

The run order the collection walks through.

Record what happened

Including the dead ends. A history of only successful fixes leaves the next engineer to retry every one of them.

Match on symptoms

Same symptoms as March, not the same as May. No precedent means no match, stated plainly.

Write the timeline live

Notes as it happens, so the postmortem is not written from memory at 3am.

See the repeats

Incidents that have happened more than once, surfaced as the ones to fix rather than mitigate.

API surface

The whole agent, endpoint by endpoint.

Endpoints

  • POST /api/incident โ€” record a resolved incident: symptoms, timings, what was tried, what worked, root cause.
  • POST /api/depends โ€” record the service graph: what it depends on and what depends on it.
  • POST /api/open โ€” open a live incident from its symptoms; matched against precedent, including what was already ruled out.
  • POST /api/note โ€” add to the timeline as the incident unfolds.
  • POST /api/resolve โ€” close it out with what worked and the root cause, back into memory.
  • GET /api/postmortem/{incidentId} โ€” the write-up, assembled from the timeline as it was written.
  • GET /api/recurring/{serviceId} โ€” what keeps happening, and so is worth fixing properly rather than mitigating again.