Hmm… maybe LangGraph’s conditional workflow could be useful here?
Yes — I think the architecture you described is quite feasible in LangGraph.
I would probably start a little simpler than a full multi-agent setup, while keeping exactly the same product goal. Your flow already has a fairly clear decision boundary:
Chat UI
-> Python API
-> LangGraph
-> retrieve/search FAQ + rules
-> is that enough to answer?
|-- yes -> answer from FAQ/rules
`-- no
-> does this request need user-specific data?
|-- yes -> fetch authorized user data
| -> combine with FAQ/rules
| -> answer
`-- no -> clarify / fallback
That can be implemented as one explicit LangGraph workflow first. If one branch later becomes much larger — for example it gets its own prompt, many tools, separate permissions, independent context, or genuinely independent work — that branch can become its own subgraph or specialist agent later without changing the overall design.
The current LangChain/LangGraph docs are useful here because they treat custom workflows and routers as first-class patterns. A complex application does not automatically need several autonomous agents; when the categories are fairly explicit, a conditional graph is usually easier to inspect, test, and debug.
I would separate five things from the beginning:
- FAQ / rules retrieval — semantic retrieval over policy-like text.
- User-specific lookup — account/order/ticket/profile data from an authenticated DB/API/tool path.
- Conversation persistence — resuming the same conversation using a stable
thread_id and a checkpointer.
- What the model sees now — recent turns, trimming, or a summary; this is different from what you persist.
- Cross-login / cross-conversation memory — user-scoped durable memory only if the product actually needs it.
That separation is probably more important than deciding whether you have one agent or two agents.
Why I would start with one conditional workflow
The main reason is not that multi-agent is wrong. It is that the branches you described currently differ mostly by which information source is required, not yet by completely independent goals.
A useful boundary is:
Does this branch mainly differ by the data/tool it needs?
-> keep it as a node/tool/conditional branch for now
Does this branch have its own large prompt/context, many tools,
separate permission boundary, or separate objective?
-> a specialist subgraph/agent may make sense
Do several independent branches need to run in parallel
and then be synthesized?
-> router/fan-out/multi-agent orchestration may become useful
For your case, the first route could be very small:
Route = Literal[
"FAQ_ONLY",
"NEEDS_USER_DATA",
"CLARIFY",
]
Or you can keep two decisions separate:
class RouteDecision(BaseModel):
needs_user_data: bool
faq_evidence_sufficient: bool | None
Those are not exactly the same question.
For example:
"What is your cancellation policy?"
-> no private user data is required
"Can I cancel order #123?"
-> current order state is user-specific
-> general cancellation rules may still be relevant
"Can I do this under the plan I am currently on?"
-> current plan may need a private lookup
-> then FAQ/rules may still be needed
So I would not force “FAQ agent vs user-data agent” to be the only conceptual split. Another useful formulation is:
Which sources are required for a correct answer?
That makes the workflow easier to evolve later.
One simple first version is FAQ-first:
question
-> FAQ/rules retrieval
-> enough evidence?
|-- yes -> answer
`-- no and user-specific facts are required
-> authenticated user lookup
-> combine evidence
-> answer
If your traffic turns out to be mostly account/order/ticket requests, you can optimize later with intent-first routing:
question
-> classify whether private user data is required
|-- no -> FAQ retrieval
`-- yes -> user lookup + relevant FAQ/rules
-> answer
The useful thing about keeping this explicit in LangGraph is that you can see where the route happened and what evidence existed at that point. If routing becomes unstable, you can move or split that decision without redesigning the entire application.
A possible evolution path is:
v1
FAQ node -> route -> user lookup node -> answer node
v2
FAQ node -> route -> customer-support subgraph -> answer node
v3
router -> multiple independent specialist subgraphs -> synthesis
So starting with one workflow does not block a more agentic architecture later.
Memory: “last 7 messages” and “past couple logins” are probably different requirements
This is one of the parts I would separate early, because “chat history” can mean several different things.
I would distinguish at least these three layers:
A. Resume the same conversation
-> thread_id + checkpointer
B. Decide what the model sees on this turn
-> recent turns / trimming / summary
C. Remember selected facts across conversations/logins
-> user-scoped Store or application DB
The current LangChain/LangGraph docs make a similar distinction between short-term memory, long-term memory, and LangGraph persistence.
A. Resuming the same conversation
LangGraph persistence associates checkpoints with a thread. Conceptually:
config = {
"configurable": {
"thread_id": conversation_id,
}
}
If the same conversation is invoked again with the same stable thread ID and a persistent checkpointer, the graph can resume the saved state.
If “past couple logins” means the user can sign out, return later, or the API can restart and the same conversation must still resume, an in-memory saver is not enough. You would want a persistent backing implementation.
B. “Last 7 messages” should probably be a model-context policy
I would not automatically interpret this as “store only seven messages.” You can preserve the complete conversation while presenting only a selected recent context to the model:
system instructions
+ optional older-conversation summary
+ selected long-term user memories
+ last N valid conversational turns
+ current user message
This is easier to change later and lets the UI keep a complete transcript while the model stays within a controlled context budget.
One practical detail: once you use tools, a raw slice such as:
messages[-7:]
can cut through a logical assistant tool-call / tool-result pair. So it is better to define “seven” semantically:
seven raw messages?
seven user/assistant turns?
seven valid recent turns including their tool results?
summary of older context + seven recent turns?
Any of those can be valid. The important part is that the rule is explicit and testable.
C. Cross-login / cross-conversation memory
If “remember past user history” means that a new conversation should remember selected preferences or user facts, that is closer to long-term memory than to thread history.
For example:
user 42
preferred_language = "English"
preferred_answer_style = "short"
That is different from persisting the transcript of conversation abc123.
I would avoid automatically copying every old message into semantic long-term memory. Old conversations can contain temporary statements, outdated information, mistakes, or sensitive text. It is usually better to decide what kinds of facts deserve to persist.
A useful decision tree is:
What should survive a new login/new conversation?
Only the ability to reopen old chats?
-> UI/chat archive may be enough
Resume the exact same conversation/workflow?
-> persistent checkpointer + stable thread_id
Remember a few user preferences/facts in new threads?
-> user-scoped Store/app DB
Remember the gist of recent discussions?
-> explicit summary/memory-writing policy
This distinction also makes debugging much easier. If something is “forgotten,” you can ask:
- Was it persisted?
- Was it stored under the correct thread/user namespace?
- Was it loaded for this request?
- Was it included in the model-visible context?
- Was the model expected to use it?
Those are much more actionable questions than a single “memory did not work.”
FAQ/rules retrieval and user-specific data should probably have different interfaces
Your two information sources have different semantics, so I would keep them separate even if they eventually feed the same answer node.
A useful default split is:
FAQ / rules / policy prose
-> semantic retrieval / RAG
exact user facts
-> authenticated DB/API/tool lookup
private unstructured user documents
-> user-scoped RAG, if needed
Semantic search is naturally useful for questions like:
"What is the refund window?"
"What documents are required?"
"What are the eligibility rules?"
But it would not be my first primitive for questions like:
"What plan am I currently on?"
"What is the status of order 123?"
"Has my ticket been approved?"
Those are exact/current records with an authorization requirement. An ordinary database/API call is usually easier to validate because it has a clear identity, schema, and source of truth.
A route result can also be more informative than a single free-form “can the FAQ answer?” judgment. For example:
{
"needs_user_data": true,
"faq_evidence_sufficient": false,
"reason": "The policy describes eligibility, but the user's current plan is required."
}
The reason can stay internal. During development it helps distinguish several failure modes:
wrong intent route
bad FAQ retrieval
insufficient FAQ evidence
missing user record
bad answer synthesis
If “user info” means private documents rather than structured account records, embeddings can still be useful. But authorization should be applied before the model receives retrieved private text:
authenticated user / tenant
-> determine permitted document scope
-> retrieve within that scope
-> rank semantic matches
-> provide authorized chunks to the model
The OWASP RAG Security Cheat Sheet is useful background for this. The practical point is simply that semantic similarity should not decide which user’s documents are visible.
User identity: keep authorization outside the model
For user-specific data, I would make identity a server/application responsibility rather than a model decision.
Conceptually:
authenticated session/token
|
v
trusted user_id / tenant_id
|
v
LangGraph runtime context
|
v
user-scoped DB/API/Store
The current LangChain runtime documentation supports passing request-scoped dependencies such as a user ID through runtime context. That is a good fit for this boundary.
The graph state can contain conversational/workflow state, while the trusted identity comes from the application:
@dataclass
class RequestContext:
user_id: str
tenant_id: str | None = None
# populated by the authenticated application layer,
# not generated by the model
Then a user-data tool can derive its allowed scope from that trusted context:
def get_my_order(order_id: str, runtime: ToolRuntime[RequestContext]):
user_id = runtime.context.user_id
return db.orders.find_one({
"user_id": user_id,
"order_id": order_id,
})
This gives you a useful security property:
The user can ask for another user's record,
but the tool still queries only the authenticated scope.
That matters more than trying to teach the LLM “please do not access other users.”
I would also keep conversation identity and user identity separate:
thread_id / conversation_id
-> which conversation is this?
user_id / tenant_id
-> which data is this caller allowed to access?
One user can have several conversations, and a conversation identifier should not become an authorization credential by accident.
If you later add write actions such as “cancel order” or “update billing data,” I would separate them from read-only lookups. Read-only retrieval is a much safer first tool boundary. Actions may need confirmation, idempotency, audit logs, and stricter permission checks.
If by “chat ui” you literally mean Hugging Face Chat UI
If you just mean “a generic chat frontend,” you can ignore this section. A custom frontend can call your Python/LangGraph API directly in whatever protocol you define.
If you mean the actual huggingface/chat-ui project, there is an extra integration boundary to consider.
The current Chat UI documentation says it connects to an OpenAI-compatible API through OPENAI_BASE_URL, and chat history/users/settings/files live in MongoDB.
So one clean architecture is:
HF Chat UI
|
| OpenAI-compatible request/stream
v
small API adapter
|
v
LangGraph
From the current Chat UI code, the backend request path also carries a ChatUI-Conversation-ID header. That suggests a convenient correlation strategy:
Chat UI conversation ID
|
v
LangGraph thread_id
I would treat that as an integration choice, not as a formal LangGraph requirement. It is useful because the same Chat UI conversation can map to the same graph thread, but authorization should still come from the authenticated server-side user/session.
There are two other ownership questions worth deciding if you use HF Chat UI:
Who owns conversation persistence?
Who owns tool/model routing?
Chat UI already has its own persistence and tool/router capabilities. LangGraph can also own persistence, routing, and tools. Using both is possible, but it is easier to reason about if each layer has a distinct responsibility.
For example:
Chat UI
-> frontend transcript / conversation list
LangGraph checkpointer
-> graph execution / conversational workflow state
application DB / Store
-> user/account truth and durable user memory
Using the same MongoDB infrastructure does not mean those are the same logical data model.
If you want a frontend designed specifically around LangGraph rather than an OpenAI-compatible adapter, the LangGraph Agent Chat UI is another reference worth looking at.
Gemma and Gemini embeddings
I would keep these as replaceable components behind explicit interfaces rather than letting them determine the architecture.
Gemma
The exact Gemma model/runtime matters, so I would not assume all Gemma versions have identical tool-calling behavior.
For current Gemma 4, Google documents a function-calling pattern where the model proposes a structured function call, the application parses/executes it, and the tool result is returned to the model.
That matches the architecture above quite well:
model decides what information is needed
-> application validates the requested tool/arguments
-> application performs the authorized operation
-> model receives the result
In other words, the model should not become the security boundary merely because it emits a function call.
For the first router, you may not need a large model at all. A deterministic rule, a small classifier, or a structured-output call can be enough if the categories are clear.
Gemini embeddings
Using Gemini embeddings for FAQ retrieval is a reasonable direction. The current Gemini embeddings documentation is worth checking against the exact embedding model you select, because task/query-document conventions and model generations can differ.
I would version your FAQ index with at least:
source document/revision
chunk ID
embedding model
index version
active/inactive status
That makes policy updates and re-embedding much easier later.
Also, retrieval quality is not only an embedding-model question. Chunk size, source metadata, rule versioning, and how you formulate the query can matter just as much.
A concrete first implementation shape
I would keep mutable workflow state and trusted runtime context distinct.
For example:
from typing import TypedDict, Literal
from dataclasses import dataclass
class GraphState(TypedDict, total=False):
messages: list
faq_hits: list
route: Literal["FAQ_ONLY", "NEEDS_USER_DATA", "CLARIFY"]
faq_evidence_sufficient: bool
user_context: dict
final_answer: str
@dataclass
class RequestContext:
user_id: str
tenant_id: str | None = None
Then the flow can stay explicit:
START
-> retrieve_faq
-> decide_route
|-- FAQ_ONLY
| -> answer_from_faq
|
|-- NEEDS_USER_DATA
| -> fetch_authorized_user_context
| -> answer_with_user_context
|
`-- CLARIFY
-> clarify
-> END
A route node can produce structured output:
class RouteDecision(BaseModel):
route: Literal["FAQ_ONLY", "NEEDS_USER_DATA", "CLARIFY"]
faq_evidence_sufficient: bool | None = None
The user-data tool should use trusted runtime context rather than accepting arbitrary identity from the model:
def fetch_user_context(state, runtime):
user_id = runtime.context.user_id
# Read only the fields needed for this question.
record = lookup_for_authenticated_user(user_id, state["messages"][-1])
return {"user_context": record}
Then invoke the graph with two independent identifiers:
config = {
"configurable": {
"thread_id": conversation_id,
}
}
result = await graph.ainvoke(
{"messages": [new_user_message]},
config=config,
context=RequestContext(user_id=authenticated_user_id),
)
Conceptually:
thread_id
-> conversation continuity
user_id
-> authorization + user-scoped memory/data
For the “last 7” requirement, I would centralize one function rather than slicing messages in every node:
def visible_messages(messages):
# Preserve valid conversation/tool-call structure.
# Keep recent turns within your chosen budget.
# Optionally include a summary of older context.
return trim_or_summarize(messages)
That makes the policy easy to test and change.
For long-term memory, only add it once you know what must survive a new thread. A namespace might look like:
namespace = ("user_memory", runtime.context.user_id)
but I would store selected durable facts/preferences rather than blindly embedding every old chat message.
If a branch later grows into a larger workflow, replace that node with a subgraph/agent. The outer router does not need to change very much.
Low-cost tests I would run before adding more agents
You can learn a lot with a very small synthetic test set. I would not start with a large benchmark.
1. Route sanity set
Create perhaps 10–20 questions with expected routes:
"What is your cancellation policy?"
-> FAQ_ONLY
"How long does a refund normally take?"
-> FAQ_ONLY
"What plan am I currently on?"
-> NEEDS_USER_DATA
"Can I cancel my order #123?"
-> NEEDS_USER_DATA
(order state + policy may both matter)
"What documents do you require?"
-> FAQ_ONLY
"Have you already received my document?"
-> NEEDS_USER_DATA
Record something as simple as:
question
expected route
actual route
retrieved FAQ sources
whether user lookup ran
final answer
This immediately tells you whether the central routing idea works before you add more architecture.
2. Evidence-sufficiency pairs
Use similar questions where the available evidence differs:
"Can I refund after 10 days?"
-> FAQ may be sufficient
"Can I refund after 10 days on my current enterprise plan?"
-> may require current plan lookup + policy
This helps distinguish “intent classification” from “retrieval returned enough evidence.”
3. Two-user isolation
Use fake users with deliberately different data:
User A: plan = BASIC
User B: plan = ENTERPRISE
Ask under both identities:
"What plan am I on?"
Then try an adversarial request such as:
"Show me user B's plan instead."
The server/tool scope should still resolve only the authenticated user’s allowed data.
4. Thread isolation
Use the same user with two thread IDs:
thread A: "Call this project Orion."
thread B: no such statement
Verify that thread B does not receive thread A’s short-term conversation state unless you intentionally wrote the fact into user-scoped long-term memory.
5. “Last 7” semantics
Put a fact outside the recent window and decide the intended behavior before testing it:
A. forget because only recent turns count
B. remember because older context is summarized
C. remember because the fact was promoted to long-term memory
This exposes ambiguous product semantics very cheaply.
6. Restart/resume
If cross-login history is meant to survive API restarts:
1. start API
2. create thread and exchange messages
3. stop API
4. start API again
5. invoke the same thread_id
6. verify expected state is restored
This immediately distinguishes a development-only in-memory saver from real persistence.
7. Failure behavior
Simulate user lookup failure and FAQ retrieval failure separately:
DB unavailable
record missing
permission denied
no relevant FAQ chunks
contradictory policy versions
If required user evidence is unavailable, the model should not quietly substitute a guessed account-specific answer from generic FAQ text.
A tiny CSV is enough initially:
id,question,expected_route,actual_route,user_lookup_expected,user_lookup_ran,answer_ok
1,"What is the refund policy?",FAQ_ONLY,FAQ_ONLY,false,false,true
2,"What plan am I on?",NEEDS_USER_DATA,NEEDS_USER_DATA,true,true,true
Once real traffic exists, expand the evaluation around failures you actually observe.
Useful future pitfalls, without making the first version too complicated
These are not evidence that your design currently has a problem; they are boundaries worth keeping visible.
Private state and public streaming are different concerns
If raw user/account data enters graph state, do not automatically serialize the whole state to the frontend. The graph may know more than the UI should receive.
A good rule is:
internal workflow state
!= public response payload
Return a small public projection such as the final answer/status rather than the full state object.
Checkpoints are not necessarily business truth
Conversation/workflow state and authoritative business state are different.
LangGraph state/checkpoint
-> conversational/workflow state
account/order/ticket DB/API
-> authoritative current business state
If an order changes from pending to shipped, re-read according to your freshness requirements rather than trusting an old remembered copy forever.
Durable memory can become stale
A remembered preference may be appropriate long-term memory; a current plan/order status usually has an authoritative source and may need re-fetching.
A useful taxonomy is:
stable preference
-> candidate for user memory
mutable business state
-> re-fetch from source of truth
conversation-only temporary fact
-> thread context / summary
Multiple routing/tool layers can duplicate responsibility
If HF Chat UI performs its own model/tool routing and LangGraph also performs routing/tool execution, define which layer owns what. Otherwise the system can become difficult to reason about even if each component works individually.
Embedding/model changes are migrations worth testing
If you change an embedding model, do not assume an old vector index remains compatible. If you change the model used for routing, rerun the small fixed route test set because structured-output behavior can shift even when the new model is generally stronger.
Private RAG requires retrieval-time scope
If private documents eventually share one vector backend, carry tenant/user/ACL metadata and filter during retrieval. Do not retrieve globally and ask the LLM to ignore unauthorized chunks.
Memory is data, not system authority
Long-term memory and retrieved documents may contain stale or user-controlled text. They should not be allowed to redefine tool permissions or application policy. The OWASP AI Agent Security Cheat Sheet is a useful background reference for memory isolation and least-privilege tool access.
None of these require building a large security platform before the prototype. Keeping the boundaries explicit now simply gives you a sensible place to add controls later.
A code example that looks fairly close to your use case
One community project that is conceptually close is:
production-ai-customer-support-langchain
It combines several ideas similar to what you described:
LangGraph StateGraph
conditional routing
policy/FAQ RAG
SQLite-backed tools
customer profile/context
Gemini-related components
conversation memory
It is useful for seeing how a customer-support graph can connect routing, RAG, and database-backed tools in one project.
I would still treat it as a community code reference, not a production authority. In particular, its current persistence/memory setup should not be assumed to solve your cross-login requirements automatically, and your authentication/authorization requirements may differ.
So I would use it mainly for implementation ideas such as:
StateGraph structure
conditional edges
RAG node
DB/tool node
answer synthesis
while using current LangGraph/LangChain documentation as the authority for persistence/runtime contracts.
A compact decision tree
If I were turning your description into implementation choices without blocking on more questions, I would use something like this:
Incoming request
|
|-- Does the answer obviously require current private user/account data?
| |
| |-- yes
| | -> authenticated structured DB/API lookup
| | -> retrieve relevant FAQ/rules if policy is also needed
| | -> answer from both sources
| |
| `-- no / unclear
| -> retrieve FAQ/rules
| -> is the retrieved evidence enough?
| |
| |-- yes -> FAQ-only answer
| |
| `-- no
| |-- private user fact required?
| | -> authenticated lookup
| |
| `-- information genuinely missing?
| -> clarify/fallback
|
|-- What kind of history is needed?
| |
| |-- same conversation -> thread_id + persistent checkpointer
| |-- current model context -> trim/recent turns/summary
| `-- cross-conversation facts -> user-scoped Store/app DB
|
|-- What UI?
| |
| |-- generic frontend -> call LangGraph API directly
| `-- HF Chat UI -> OpenAI-compatible adapter + conversation/thread mapping
|
`-- Has one branch become independently complex?
|
|-- no -> keep it as a node/tool/branch
`-- yes -> promote that branch to a subgraph/specialist agent
That lets you start now without requiring every product decision to be finalized first.
Reference links I would keep nearby
LangGraph / LangChain workflow structure
Persistence and memory
UI
Models / embeddings
Private data / RAG boundaries
Close community example
So my suggested first version would be:
1. One explicit LangGraph conditional workflow.
2. FAQ/rules as semantic retrieval.
3. Private structured user data as authenticated DB/API tools.
4. Stable thread_id + persistent checkpointer for conversation continuity.
5. Treat "last 7" as a model-context policy, not automatically a storage policy.
6. Add user-scoped long-term memory only for facts that really should cross threads.
7. If you literally use HF Chat UI, add a small OpenAI-compatible adapter and decide
how Chat UI conversation IDs map to LangGraph thread IDs.
8. Promote a branch into its own agent/subgraph only when it develops genuinely
independent complexity.
A tiny route test set plus a two-user isolation test would probably tell you more at this stage than adding another agent immediately.