Conversational Cross-Channel Booking Agent
A booking agent that keeps one shared inbox across WhatsApp, Instagram, Outlook, and Google instead of four separate ones.
The Setup
A solo service-business owner gets booking requests from four places that do not talk to each other: WhatsApp, Instagram DMs, email, and calendar invites. Each channel holds its own version of what is booked. Nothing reconciles them, so there is no single answer to what is confirmed, what is pending, or what got promised twice.
Why it costs real money
A solo operator handling 50-plus booking messages a day across four channels loses two-plus hours daily to triage and scheduling, time that would otherwise go to clients. Double bookings and missed DMs are not small annoyances for someone with no staff to absorb them.
Why it is not a chatbot problem
Each platform has its own data model, its own webhook delivery guarantees, and its own OAuth flow. Holding conversation state across all four, while handling duplicate messages, retried webhooks, and matching the same customer across platforms, sits in distributed systems territory far more than it sits in AI.
Where It Drifts
The obvious build is one bot per platform, each with its own database, reconciled by hand into a shared calendar. It fails in three specific places, and all three are state problems rather than model problems.
Each bot only knows the bookings it made itself. Reconciling that into one calendar is a manual step, and it is the first thing to slip once things get busy.
One FastAPI gateway and one LangGraph agent write every booking to a single Cloud SQL store. There is no second calendar to reconcile.
No idempotency check anywhere. WhatsApp, Instagram, and Outlook all retry webhook delivery, so the same event can trigger a second booking attempt.
Idempotency keys (message_id plus platform) are checked against Memorystore before a Cloud Task runs, so duplicate deliveries get rejected before they reach the booking flow.
Customer identity is scoped per platform, with no matching against phone or email. A customer who DMs on Instagram and later messages on WhatsApp becomes two unrelated records.
Not solved in v1. This is the redesign flagged as highest priority: build identity resolution before any bot logic, not after.
Path Of One Message
The shape worth noticing is the gap between the webhook and the agent. Everything else follows from refusing to do booking work inside a request that has three seconds to live.
Engineering Decisions
- 01
LangGraph's cyclic StateGraph instead of a linear chain
Booking needs confirmation loops: the agent proposes a slot, the customer declines, the agent checks availability again. A linear chain cannot re-enter a node once it has run. StateGraph allows that explicitly, so the confirmation node just routes back to the availability node on a rejection.
- 02
Cloud Tasks instead of a self-managed queue
A webhook has to return HTTP 200 within about 3 seconds or the platform retries it. The LangGraph agent can take 8 to 15 seconds to finish a booking loop. Cloud Tasks separates the two: the webhook returns immediately, and the booking work runs asynchronously on Cloud Run, with no broker or worker fleet to run ourselves.
- 03
Google Cloud Storage instead of platform CDN links
Instagram's media URLs expire in 24 to 48 hours; WhatsApp's expire after being accessed once. GCS holds the files permanently, and Cloud Scheduler runs a periodic job that refreshes profile pictures before their source CDN link expires.
- 04
Per-platform idempotency keys in Memorystore
Every major messaging platform can and will deliver the same webhook event more than once. Keys built from the message ID and platform type get checked against Memorystore before a Cloud Task is allowed to run, so a duplicate delivery never turns into a duplicate booking.
Cloud Tasks and Cloud Run push operational weight onto GCP instead of a self-run broker and worker pool. GCS costs more per request than self-hosting object storage, but it sidesteps the expiring-URL problem entirely. LangGraph's cyclic state model adds serialization overhead a simple prompt loop would not need, which is worth it only because the confirmation flow genuinely has to re-enter a node.
By The Numbers
Results
What I would measure next
Booking completion rate, from intent to confirmed slot. Webhook delivery latency, from platform send to Cloud Task execution. And how often the system actually catches a calendar conflict before it becomes a double booking, which is the number that would justify the whole design.
Stack
LangGraph
Booking confirmation needs to loop: propose, decline, re-check, propose again. StateGraph supports that natively and keeps conversation state explicit across tool calls.
Gemini
The model behind the agent's tool-calling loop, wrapped in a small LangChain-compatible class instead of called directly, so the graph can swap models without touching node logic.
MCP (Model Context Protocol)
Appointment and slot-availability tools run as their own MCP servers rather than living as inline LangChain tool functions, so the booking logic stays addressable on its own.
Langfuse
Traces every tool call inside a booking loop. Without it, a bad booking is just a support ticket with no way to see which tool ran or why.
FastAPI
Handles webhook traffic from four platforms concurrently. A synchronous framework would queue these one after another and risk webhook timeouts.
SQLAlchemy (async) + Alembic
Async ORM for conversation state, customers, and appointments, with Alembic tracking schema changes as the booking model grew.
Cloud Run
Scales to zero between bursts of booking traffic and scales out during a busy DM window, without servers to manage directly.
Cloud Tasks
Separates webhook receipt from the actual booking work. The webhook needs an immediate 200; the agent needs 8 to 15 seconds.
Cloud SQL
Relational storage for conversations, customers, and appointments, reached through the Cloud SQL IAM connector instead of a static password.
Memorystore for Redis
Sub-millisecond reads for idempotency keys and short-lived session state, managed and Redis-compatible so the client code does not change.
Cloud Scheduler
Runs the media-refresh job on a fixed cadence, catching CDN links before they expire without a cron daemon to babysit.
Google Cloud Storage
Permanent home for media whose platform CDN links expire, refreshed by the Cloud Scheduler job.
WhatsApp Business Cloud API
The primary channel for the target market. The Cloud API skips the server management that On-Premises would require.
Microsoft Graph (Outlook)
Corporate clients live on Outlook. Graph API gives calendar and email access under one OAuth flow instead of two separate integrations.
Firebase Cloud Messaging
Push notifications for booking confirmations on both iOS and Android from one integration.
What I'd Rebuild First
- 01
Cross-platform identity resolution, before any bot logic. Matching a WhatsApp number to an Instagram handle to an email address is the hardest unsolved part of the system, and it got treated as an afterthought instead of the foundation.
- 02
Correlation IDs in logging from the first commit. A booking failure that crosses a WhatsApp webhook, a Cloud Task, a LangGraph node, and a Calendar API call needs one trace connecting all four. Retrofitting that is painful.
What The Integrations Taught Me
- —
Multi-platform OAuth is not one problem, it is four. Each platform expires tokens differently, scopes refreshes differently, and re-subscribes webhooks differently. The auth layer ended up being more code than the agent itself.
- —
Idempotency is rarely right on the first try. The natural key for 'already processed' differs by platform, and getting that schema right took more than one pass.
- —
Requiring every piece of state to serialize to Redis forces early discipline about what belongs in the state graph. It ruled out a few designs that would have been painful to scale later, and in hindsight that was the point.
Solving a similar problem?
I'm open to conversations about production AI systems -agentic workflows, RAG pipelines, or messy integration problems like this one.