
All right, so what you just saw wasn't some futuristic demo reel. It was a live interaction with an AI avatar that can actually take your order, talk about your business, and do real work on your website. I built that restaurant ordering system to show exactly what's possible right now, not in six months. And the best part? You don't need a PhD in machine learning to wire this up yourself.
The whole thing runs on an application called Synthesia, hooked into a lightweight real-time communication layer and whatever large language model you want driving the brain. I am going to walk you through exactly how I built it, from choosing the avatar's face to making it actually place items in a cart. No fluff, just the steps.
Getting your AI avatar ready inside Synthesia
First things first, head to synthesia.io and log in. Once you are inside the platform, you want to go straight to the Avatars section. This is where you build the face your customers will actually talk to.
Synthesia gives you a few different paths here. You can pick from existing avatars, either animated or a real person. You can also generate a custom 3D avatar, which is what I did for this demo. You choose the base avatar, set the pose, pick the space it stands in. You can even upload your own image for the background space if you want something branded. Different outfits are available too.
There is also an AI generation option where you just describe what you want and it spits out versions of an avatar for you. Once you have something you like, you hit save and it lives in your workspace ready to be deployed.
The thing to remember here is this avatar is going to be the literal face of your business for anyone who lands on your site. So don't just grab the first cartoon you see. Think about whether it matches the tone of what you are selling. A crisp, professional-looking avatar works for a restaurant or a consultancy. Something more playful might work for an e-commerce store selling gadgets.
Connecting your avatar to a real web application
This is where most people get stuck, thinking the avatar is just a video file you embed. It is not. It is an interactive agent that listens and responds in real time. To get that working on your own website, you need to head to the Developer section inside your Synthesia workspace.
Here you will find a quick-start guide with actual code. The pattern is simple: you import the Synthesia package, reference your specific AI avatar ID, initialize the session, and the avatar is live on your page. The platform also links out to a full repository that shows exactly how to spin up an interactive avatar application.
So how does the conversation actually flow under the hood?
Your web application talks to the avatar. When you speak, your audio gets sent to a service called LiveKit. LiveKit takes that speech, converts it to text, and hands that text off to your AI agent. That agent is where you plug in your knowledge base, your RAG systems, your database, everything. The agent can also call tools, like booking a meeting or placing an order, which is where things get powerful. Once the agent formulates a response, it passes that output back to Synthesia to generate the avatar's voice and facial movements. LiveKit ships that final result back to your web application. You see the avatar speak.
What I did next was take the entire repository link and hand it over to Claude and GPT. I gave the model the full prompt, pointed it at my tech stack, and had it read through the official LiveKit documentation to figure out all the environment variables we would need. I also pointed it at my own knowledge base, a single file containing everything about my YouTube channel and my Skool community. That became the brain the avatar would use to answer questions. The full prompt I used is in the description if you want to grab it.
Setting up the keys that make everything talk
You need four critical pieces to make this work. GPT built the application, but I had to feed it the right credentials.
Finding your avatar ID and voice ID
Back in your Synthesia workspace, when you select an avatar, you can copy its ID directly. Same for the voice. You click the voice, copy its ID, and pass both of those to your application so it knows exactly which face and which voice to load. I copied my avatar ID and pasted it into the ENV file configuration GPT generated for me.
LiveKit Cloud credentials
LiveKit Cloud is completely free to start. You sign up, create a project, and navigate to Settings then API Keys. Click Create API Key, give it a name like "demo," and it generates your URL, your API key, your secret, and the full environment variable block. You copy that block and drop it straight into your application's ENV file. I will delete my keys after this tutorial, but the process takes thirty seconds.
Synthesia API key
In the Synthesia developer section, there is a dedicated API Keys tab. Click Create API Key, give it a name, and for the interactive avatar you need to grant access. Set the expiry to whatever makes sense. Seven days works for testing. Once created, copy the key and pass it to your application as the Synthesia environment variable.
OpenAI key
You can use any model provider here, OpenRouter works too, but for this demo I stuck with OpenAI. I generated a key, copied it, and dropped it into the ENV as well. Now the agent has a brain.
After all the keys are set and the knowledge base file is wired in, the application is ready to test.
Watching the AI avatar actually answer questions
I clicked the web application preview. The avatar loaded on the right side of the screen and I hit start. After a brief pause, it spoke.
It introduced itself as my AI assistant and asked what I was curious about. I asked it who I am and how many subscribers I have. It correctly identified that I am a former AI and software engineer at Amazon and Microsoft, now teaching practical AI development through YouTube and running my AI Builders community on Skool. It didn't hallucinate a subscriber count. It told me honestly it didn't have the real-time number and suggested checking YouTube directly. That is exactly what a good agent should do: know what it knows and not invent data.
I then asked what is inside my Skool community. It listed the topics correctly: building AI agents, setting up automations with n8n, organizing second brains, using Claude Code for applications. It mentioned the learning roadmap, templates, and the Q&A space. It then asked me what kind of project I was working on. This is the conversational loop you want for customer service: answer, then re-engage with a follow-up question.
I also tested the mute and end-call functionality. Both worked. This is a fully interactive session, not a pre-recorded clip.
How the restaurant ordering system really works behind the scenes
Now the intro demo, the one with the lacquered duck and the dim sum, is a different beast. That avatar is not just answering questions. It is navigating a menu, making recommendations, and adding items to an order total with tax calculated in real time. So let me break down the actual architecture I drew up for that system.
A guest lands on your restaurant's web application. The application connects to the LiveKit Cloud immediately. Right now, for simplicity, I am using browser storage to hold the order and the cart state. If you push this to production, you would swap that out for a proper database. The communication protocol between your web app and the server flows entirely through LiveKit.
Here is the loop step by step:
- The guest speaks their order ("I want the lacquered duck for six people").
- LiveKit streams that audio to a speech-to-text service, in this case Cartesia, which transcribes it into text.
- That text gets fed to OpenAI or whatever model you have wired in.
- The model processes the request against your menu knowledge base and the current cart state, then generates a response and, critically, performs an action, like adding a whole duck to the order and updating the running total.
- I also built in a logging system that writes the conversation to a local MD file so you can review interactions later and fine-tune the model's behavior.
- The model's text response goes to Synthesia, which generates the avatar movements and the speech.
- LiveKit serves that response back to the web application.
- The guest sees the avatar confirm the order and sees their bill updated on screen.
What makes this different from a simple chatbot is the tool calling. The avatar does not just say "I recommend the duck." It says that, then it hears "yeah let's do it," then it actually places the order. The customer sees the line item appear. The total updates. The system knows a whole duck serves six and doesn't ask redundant follow-ups unless it needs clarification. That is the difference between a gimmick and something you could run a business on.
The full code for that restaurant application is linked in the description. Download it, tear it apart, and see how you can slot it into whatever you are already running.
Points clés à retenir
- Un avatar IA interactif peut être intégré à n'importe quel site web, pas seulement pour répondre à des questions, mais pour exécuter des actions réelles comme passer une commande ou réserver un rendez-vous.
- Synthesia fournit les visages et les voix, et une section développeur complète avec le code nécessaire pour embarquer l'avatar dans votre application.
- LiveKit Cloud sert de couche de communication en temps réel entre le navigateur du client, le modèle de langage, et le service de génération d'avatar. C'est gratuit pour démarrer.
- Le système fonctionne en boucle: la voix du client est transcrite en texte, envoyée au LLM avec votre base de connaissances, le LLM peut appeler des outils (ajouter un article au panier, naviguer sur le site), la réponse est convertie en parole et en mouvements faciaux par Synthesia, puis renvoyée au client.
- Vous pouvez connecter n'importe quel modèle de langage (OpenAI, Open Router, etc.) et lui fournir votre propre base de connaissances pour que l'avatar parle précisément de vos produits, vos services, ou votre contenu.
- Un système de logging intégré enregistre les conversations dans un fichier local, ce qui vous permet d'affiner le comportement de l'agent avec le temps.
You have the pieces now. The avatar, the real-time pipeline, the tool execution. The restaurant demo is not a toy. It is a blueprint. Whether you run an e-commerce store, a consultancy, a SaaS landing page, you can bolt this exact same pattern onto what you already have and give your visitors someone who actually does something instead of just filling out a contact form. That changes how fast deals close.
