Order intake agent
About this agent
The first of three unbundled retail agents that can be adopted on their own or chained together. This one turns a buyer's WhatsApp voice note or shorthand text โ "bhai 5 bora chini and same aatta as last time" โ into a clean, structured order the rest of the pipeline can act on. Audio goes through a pluggable speech-to-text provider first; the resulting transcript is then parsed by an LLM call that is grounded in that merchant's own item-name aliases and order patterns, pulled from long-term project-scope memory, so "chini" resolves to "Sugar" for this shop specifically rather than a generic dictionary. If the model's output does not parse into valid JSON, or yields no items, the order is flagged for human review rather than guessed at โ no item name, quantity or unit is ever invented to fill a gap. It does not check stock, print packing slips or touch money; that is deliberately left to the Warehouse & Dispatch and Ledger & Recovery agents downstream, so a shop can take just the time savings here without exposing inventory or ledger data to the same pipeline.
What changed with Zenmem?
The same agent, built twice against the same contract โ once on Zenmem, once on MongoDB + LangChain/LangGraph.
Before โ after
Before With Zenmem
What the team gained
- The parse call reads the merchant's own alias scope and writes a summary back, so the alias list improves as orders come in with no training loop to run.
- Aliases and order patterns are keyed by merchant, so one shop's buyers can never influence another shop's parse.
- The raw transcript and the parsed order are session-scoped as an audit trail the parser never reads back from, keeping a bad parse from teaching itself.
- A new item alias is a document, not a row in a mapping table someone maintains.
- Passing memory into the parse and saving the summary out are parameters on the same call, so there is no field-by-field mapping to keep in step.
How memory is scoped
Project scope, keyed by merchant id, holds the durable stuff: item-name aliases and order patterns that should carry over from order to order but must never leak from one shop's buyers into another's. The LLM parse call reads this scope with passMemory=True and writes a summary back with saveInMemory=True, so the merchant's alias list improves as more orders come in. Session scope, keyed by the WhatsApp thread, holds the raw transcript and the parsed order for this one order only โ an audit trail, not a source the parser reads from.
How it works
The take_order() pipeline, start to finish.
Transcribe
A voice note goes through the pluggable speech-to-text provider (mock by default, OpenAI Whisper in production); a text message is used as-is.
Pull merchant aliases
Fetches this merchant's top item-name aliases and order patterns from project-scope memory before parsing anything.
Parse with grounding
An LLM call turns the transcript into strict JSON, using the merchant context to resolve shorthand, then saves a summary of the exchange back into that memory.
Validate or flag
The JSON is turned into typed OrderItem records; a malformed response or zero parsed items flags the order needs_human_review instead of guessing.
Log to session
The transcript and the parsed order are written to this order's session memory as an audit trail.
What it does
The agent's public surface, as a Python class.
Capabilities
- take_order(...) accepts either a text_message or an audio_path and requires at least one of the two.
- Transcribes voice notes via a pluggable SpeechToTextProvider โ a mock provider offline, OpenAI Whisper when RETAIL_AGENTS_STT_PROVIDER=whisper.
- Parses free-form or shorthand transcripts into strict JSON via callLLM, grounded in the merchant's own item-alias memory.
- Produces typed Order and OrderItem records: raw_text, normalized item_name, quantity, unit and optional notes.
- Flags an order needs_human_review when the model's JSON is malformed or contains no items, rather than inventing a result.
- Falls back to a regex-based heuristic order parser in mock mode, so the pipeline runs with no LLM credentials.
- Writes the transcript and parse outcome to the order's own session memory as an audit trail.