Skip to main content
KANG LAB.
Project Index
PROJECT / 002Active

AI Xiaoer

A private-room dining advisor — guests scan the table code and talk to an in-store AI for personalised dish recommendations, while the owner controls the menu, availability and recommendation weights from a back office

Year
2026
Status
Active
Categories
ai / web / python
Updated
2026.09.30
  • Taro
  • React
  • FastAPI
  • LangGraph
  • DeepSeek
  • MySQL
  • Redis
AI Xiaoer — visual

Overview

AI Xiaoer is a dining advisor built for the private rooms of a single restaurant. Guests scan the code on their table sticker, land on a page owned by that restaurant, and describe their party size, budget and tastes in natural language. The AI answers with dishes drawn from the restaurant's real menu, each with a stated reason. The owner maintains dishes, availability and recommendation weights from a back office, and so decides how the AI recommends.

The scope is deliberately narrow: recommendation only — no ordering, payment or kitchen integration. There is no user account system; scanning is enough, and taste preferences are remembered only within a single session.

Problem

Three concrete pain points in private-room dining, judged against this restaurant and comparable Chinese restaurants:

  • The host carries a heavy decision load — eight people, differing tastes, dietary constraints to respect, a budget to hold, and dishes that need to look impressive. Twenty minutes with the menu is often not enough to decide;
  • Guests hesitate to keep asking — they want to know whether a dish is spicy or how large the portion is, but stop asking after two or three tries and order by guesswork;
  • The owner cannot push what they want to sell — new dishes get no cold start, high-margin dishes sit on page three, and staff turnover is high: experienced waiters sell, new ones cannot memorise the menu.

Almost every AI ordering product on the market is fighting for a platform traffic entrance — cross-restaurant recommendations that serve the platform's interest. None of them is aimed at this restaurant's own private rooms. So the project takes a single-point validation approach: private-room recommendation only, with weights configured directly by the owner. The LLM defaults to a cloud API to keep cost down, while the architecture retains a local-deployment fallback.

Architecture

Guest client  Taro 4 + React (H5 first, mini program as second compile target)
   ↓  table code / short code / native mini program code
Backend       FastAPI + LangGraph
   ├─ candidate recall (on sale + achievable spiciness match)
   ├─ LLM gateway (OpenAI-compatible: DeepSeek by default, switchable to Qwen / Zhipu / Kimi / vLLM / Ollama)
   └─ output validation with fallback
   ↓
Data          MySQL 8 + Redis 7
   ↓
Back office   React 18 + Vite + TDesign

Deployment is a single all-cloud host: Nginx (static assets, reverse proxy, HTTPS) with uvicorn, MySQL and Redis. No GPU is needed on the application server.

Features

Guest client

  • Scan to enter and converse with the AI in natural language; replies stream token by token over SSE;
  • Recommendation cards carry four distinct states: addable, added, sold out today, and conflicting with this table's dietary constraints (not recommended, but the reason is stated plainly);
  • Guests without a table code get a takeaway entry point; after a table is closed the guest is no longer left retrying; dropped messages on a weak connection can be resent;
  • Past recommendation cards collapse for review, and the in-session board persists.

Back office

  • Dish management: create, edit, delete, sortable columns, image upload (resized and re-encoded to WebP server-side), category management;
  • Excel batch import: tolerant header-name matching row by row, with same-named dishes updated in place and a one-click duplicate merge for existing data;
  • Availability switch: marking a dish sold out stops it being recommended in every session immediately, and the AI answers honestly with an alternative when asked directly;
  • Recommendation strategy: AI persona and bulk dish weight configuration;
  • Private room management: short-code generation, table stickers and native mini program codes, downloadable as one package;
  • Session management: a list and detail view for the day's sessions (full transcript plus recommendation log), with batch termination by room and undo;
  • Dashboard: overview, 7-day trend and dish ranking.

Design notes

Design decisions hardened after running into each of these:

  • Anti-hallucination is three gates: recommendations must come from the on-sale candidate set; factual fields such as price and spiciness are always read from the database; and when model output fails validation the system degrades rather than rendering unvalidated text to the guest;
  • Session timeouts are two separate parameters: a quiet threshold (30 minutes by default, which only affects how the back office labels a session) and a resumability window (180 minutes by default, which decides whether a guest can return to their session). Sharing one value originally meant a banquet longer than three hours lost its history the moment nobody talked to the AI for a while;
  • The business-day boundary is 4 a.m. and is not hard-coded: the boundary, the quiet threshold and the resumability window are all configurable from the back office and take effect immediately;
  • Spiciness is split into a default level and an achievable set: when a guest says they avoid chilli, dishes that can be made mild stay in the running, instead of every chilli-bearing dish being excluded;
  • One source of truth for LLM configuration: back-office values override environment variables, keys live server-side only and are masked in the UI. The seed export excludes the whole llm_* group, so a half-applied override — base URL and model set, key missing — can no longer silently shadow a valid key in the environment;
  • Tests never depend on external services: the backend runs pytest against an in-memory SQLite database, requiring no MySQL, Redis or network, and every test that would call the LLM mocks the request instead.

Status

Across 9 working days from 2026-09-11 to 2026-09-24 and 176 commits, the guest conversation loop, six back-office pages and the production deployment are in place. Daily logical database backups and an isolated restore drill have been verified on the server.

Roadmap

  • Shared sessions per room, with a common shortlist and shared dietary constraints;
  • WeChat authorisation and a long-term taste preference profile;
  • Structured dietary hard filters, ordered-dish awareness and time-of-day strategy;
  • Off-host backup to object storage, removing the single-disk failure risk.

Privately deployed for one restaurant, running on its own server; the guest client is distributed through table stickers and is not publicly indexed.

Changelog

  1. 2026.09.11
    • Repository initialised; deployment settled as a single all-cloud host, with the guest client shipping as Taro H5 first and the mini program as a second compile target
    • Product plan raised to v0.3: back-office session management added, session lifecycle and recovery rules made explicit
    • Backend skeleton and the guest-side session-creation path landed; LLM gateway introduced (provider abstraction over DeepSeek, switchable to a local vLLM instance)
    • Core recommendation chain wired end to end: candidate recall and scoring, prompt assembly, output validation with fallback
    • Guest conversation loop completed: intent routing, constraint extraction, streaming send-message endpoint
  2. 2026.09.12
    • Back-office APIs completed: dishes, private rooms, recommendation strategy, session management, system settings
    • Excel batch import shipped with template download; re-importing a same-named dish updates it, plus a one-click duplicate merge
    • Spiciness split into a default level and an achievable set, so recall decides on whether a dish can be made to the guest's requested level
    • Admin shell built out (Vite + React + TDesign) behind login authentication
  3. 2026.09.14
    • Five admin pages completed (dishes, private rooms, recommendation strategy, sessions, settings), lists unified into an adaptive card layout
    • Business-day auto-archiving scheduled task and Alembic database migrations added
    • Guest client switched to the real backend, replacing every mock; fixed the single root cause behind non-streaming replies and missing cards
    • Past recommendation cards collapse for review; the in-session board and shortlist persist
  4. 2026.09.15
    • Event tracking wired up, with card_click and summary_view reported from the guest client
    • Dish image upload added, uniformly resized and re-encoded to WebP with EXIF stripped; cards reworked to image-left, text-right
    • Fixed lost session memory: party size and dietary constraints no longer dropped across turns
    • Long-conversation benchmark added: over 30 turns the cost sits in the cards, not the text
    • Single-host Alibaba Cloud deployment scripts took shape, with the back office served under /admin/
  5. 2026.09.21
    • Table stickers redesigned; scoped recommendations and a separate owner prompt added
    • Daily database backup and an isolated restore drill put in place
    • Release gate hardened: a failed dependency install, migration or health check aborts the deploy
  6. 2026.09.22
    • Production backup installation and the isolated restore drill verified
    • In-store location gate added for guests, with the service area configured from the back office
    • One-shot import of a menu package complete with images
  7. 2026.09.23
    • HTTPS enabled on the domain; release gate extended to require a 401 from unauthenticated APIs
    • Mini program adaptation completed: requests, streaming conversation, images, location and code scanning
    • Table QR codes replaced with native mini program codes so WeChat scan-to-open enters the mini program directly; the back office can package all stickers in one download
    • Added a takeaway entry point so guests without a table code do not have to go looking for one
  8. 2026.09.24
    • GCJ-02 coordinate system unified across the whole chain, fixing in-store users being judged as having left
    • Service area changed from a single point to multiple locations, with card-based editing and map verification in the back office
    • Guest shortlist additions persisted, so the back office can review whether recommendations were actually used
    • Statistics groundwork laid (indexable events, action semantics, one shared definition), and a back-office dashboard shipped: overview, 7-day trend and dish ranking
    • Mini program avatar produced; table stickers now carry the in-store Wi-Fi details

Related Projects