Test case builder
About this agent
Name a feature or module in plain words and get back a test suite traced to the code it was read from. Generation runs across seven facets, narrowed to five in smoke mode and widened in exhaustive; the output can be restricted to particular case types โ negative, boundary and the rest. A suite can be refined in place, and case ids stay stable so anything already linked to a ticket or a test run does not renumber. Coverage the indexed code cannot support lands in `coverageGaps` rather than becoming a fabricated case, and every route refuses with 409 until the desktop utility has actually indexed the workspace โ a suite generated against an empty knowledge base would look plausible while describing code nobody read.
What changed with Zenmem?
The same agent, built twice against the same contract โ once on Zenmem, once on MongoDB + LangChain/LangGraph.
Before โ after
Before With Zenmem
What the team gained
- The workspace knowledge base is the retrieval layer, so tracing a case back to the code it was read from needs no separate embedding pipeline to maintain.
- A refinement thread is a session id, which is what lets a follow-up resolve 'those' without the client restating the earlier answer.
- Session and workspace scopes stay apart, so a refinement conversation never contaminates the indexed code every suite traces to.
- Coverage the code cannot support falls out as an empty retrieval rather than a fabricated case.
- Seven facets at standard depth are seven retrievals against one store, not seven query paths to keep in step.
How memory is scoped
Two scopes. The workspace knowledge base is long-lived: the codebase the desktop utility indexed, which every generation retrieves against and which every case traces back to. A session is the refinement thread on top of it โ generate or chat allocates one, follow-ups against the same id resolve without restating the earlier answer, and closing it is terminal, so the next request takes a fresh id rather than reopening a closed one.
How it works
The run order the collection walks through.
Index first
The desktop utility builds the knowledge base. Until it has, every generation route refuses with 409 rather than inventing coverage.
Generate across facets
Seven facets at standard depth, five at smoke, wider retrieval at exhaustive.
Refine without renumbering
Case ids stay stable, so links to tickets and test runs survive a refinement.
Export where it is used
Markdown for the pull request, Gherkin for the runner, CSV for the test management sheet.
API surface
The whole agent, endpoint by endpoint.
Endpoints
- GET /health โ liveness, plus whether Zenmem itself is reachable; `degraded` means the service is up but the deployment is not.
- GET /api/v1/workspaces/{workspaceId}/status โ is the knowledge base ready, and which files were indexed. Always 200, so 'not indexed yet' is never confused with 'service is broken'.
- POST /api/v1/testcases/generate โ the main route: a feature in plain words, a depth of smoke, standard or exhaustive, and an optional type filter.
- POST /api/v1/testcases/refine โ adjust a suite in place with stable case ids.
- GET /api/v1/suites/{suiteId} โ one suite in full.
- GET /api/v1/workspaces/{workspaceId}/suites โ every suite for a workspace.
- GET /api/v1/suites/{suiteId}/export โ Markdown for review, Gherkin for a BDD runner, CSV for spreadsheet test management.
- POST /api/v1/chat โ ask about the codebase or its coverage without generating a suite.
- POST /api/v1/sessions/{sessionId}/close โ end a refinement thread. Session ids are single-use and terminal.