One of the questions AI toy brands ask most in 2026:How do I connect my AI module to our own LLM?This article sets out how the Nablai LX Series handles it in practice, and the typical architecture behind it.
1. Why you need an LLM of your own
Most toy brands start by wiring up a public cloud API (Doubao, Tongyi, ERNIE, GPT, Claude and so on) to get an MVP running. Once it works, they run into three problems fast:
- User data ownership: the public API vendor gets hold of children's conversation data, and the brand has no visibility into what happens to it
- Character consistency: a public LLM has no hard constraints on the tone, values or content boundaries of an IP character, so it breaks character easily
- Long-term cost: billed per call, a fleet of 100,000 devices can produce a monthly bill in the millions, with no cap in sight
The only way out of all three:An LLM the brand owns. It can be a private deployment, or a dedicated model running on your own private cloud.
2. How the LX Series connects
The Nablai LX Series abstracts three layers at the bottom of the stack so brands can switch LLMs freely:
| Layer | Role | Options |
|---|---|---|
| On-device ASR/TTS | Speech recognition / speech synthesis | Local on-device model (default) / cloud ASR (optional) |
| Conversation brain | Understanding + generation + character | The brand's own LLM / public API (fallback) |
| Content operations | Story library, character settings, content moderation | The brand's own content operations console / the Nablai AMS console |
Brands can leave on-device ASR/TTS exactly as it is (on-device wake < 300ms,弱网可用),只把"对话大脑"换成自有大模型,Overall response latency still stays within 1.5 seconds。
3. A four-step practical guide
Step 1: Align on requirements (1-2 weeks)
- The brand's IT team aligns with the Nablai technical team on: how their own LLM will be deployed (private server / private cloud / VPC)
- Character settings: the IP character's tone, values, conversation boundaries and how it handles sensitive topics
- End-to-end SLA: response latency, availability and what triggers content moderation
Step 2: Private deployment (2-4 weeks)
- The brand provides the LLM inference service (OpenAI-compatible API / custom protocol)
- The Nablai module firmware is updated over OTA to a specified version and the private brain mode is enabled
- Configure in the AMS console: API endpoint, authentication method and degradation policy
Step 3: Joint debugging (2-3 weeks)
- Long-tail scenario testing: weak networks, packet loss, high concurrency, malformed input
- Content moderation triggers: boundaries around self-harm, violence and politically sensitive topics
- Character consistency testing: aligning tone and values across 100 representative conversations
Step 4: Staged rollout (4-8 weeks)
- 10% of devices: monitor response latency, error rate and user activity
- 50% of devices: monitor subscription conversion and the effect of content updates
- 100% rollout: full OTA firmware push, with a seamless switch for users
4. A typical architecture
[Toy-side LX002S module] On-device ASR (wake < 300ms)
端侧 TTS(合成)
↓
HTTPS / WebSocket
↓
[品牌方私有云 / VPC]
自有大模型推理服务(OpenAI 兼容 API)
内容审核模块
角色库 / 故事库
↓
[梯度算子 AMS 后台]
设备激活 / OTA / 用户活跃度
订阅与计费(品牌方独立账户)
5. Four common questions
Q1: Is a private LLM expensive?
If you deploy an open-source model (Qwen, DeepSeek, Llama and others) privately, inference costs about ¥0.0001-0.001 per call at up to 100,000 QPS, one to two orders of magnitude cheaper than a public API. The one-time investment is GPU servers, roughly ¥300,000-1,000,000.
Q2: Will response latency get worse?
If your own LLM runs in a VPC in the same region, end-to-end latency stays within 1.5 seconds. Cross-region, it can reach 2-3 seconds. We recommend same-region deployment plus inference optimization (quantization, KV cache, prefix cache).
Q3: How is firmware OTA managed?
The Nablai AMS console supports pushing OTA independently for each brand. The module itself has staged rollout capability, so updates can be targeted precisely by device serial number, S/N range or user tag. A failed OTA rolls back automatically.
Q4: Who owns the subscription revenue?
The brand. Nablai takes no share of subscription revenue and charges only for the module BOM and an operations and maintenance fee. The complete brand-owned account system, subscriptions and content operations all sit with the brand.
6. When an LLM of your own is the wrong call
- You only want one or two AI SKUs to test the market - running an MVP on a public API is faster
- The brand's IT team cannot operate an LLM - strengthen the team first, or the private deployment becomes long-term debt
- The target user base is small (< 5 万)——私有化投入回收周期太长
If what you are aiming for is long-term IP assets plus recurring subscriptions plus brand sovereignty, the LX Series with your own LLM is the safest combination in 2026. Full analysis: https://www.nablai.com.cn/articles/lx-series-brand-sovereignty.html