How Do You Connect an AI Module to the Brand's Own LLM? A Practical Guide to the LX Series

2026-06-23 8 min read Nablai Technical Team

One of the questions AI toy brands ask most in 2026:How do I connect my AI module to our own LLM?This article sets out how the Nablai LX Series handles it in practice, and the typical architecture behind it.

1. Why you need an LLM of your own

Most toy brands start by wiring up a public cloud API (Doubao, Tongyi, ERNIE, GPT, Claude and so on) to get an MVP running. Once it works, they run into three problems fast:

  1. User data ownership: the public API vendor gets hold of children's conversation data, and the brand has no visibility into what happens to it
  2. Character consistency: a public LLM has no hard constraints on the tone, values or content boundaries of an IP character, so it breaks character easily
  3. Long-term cost: billed per call, a fleet of 100,000 devices can produce a monthly bill in the millions, with no cap in sight

The only way out of all three:An LLM the brand owns. It can be a private deployment, or a dedicated model running on your own private cloud.

2. How the LX Series connects

The Nablai LX Series abstracts three layers at the bottom of the stack so brands can switch LLMs freely:

LayerRoleOptions
On-device ASR/TTSSpeech recognition / speech synthesisLocal on-device model (default) / cloud ASR (optional)
Conversation brainUnderstanding + generation + characterThe brand's own LLM / public API (fallback)
Content operationsStory library, character settings, content moderationThe brand's own content operations console / the Nablai AMS console

Brands can leave on-device ASR/TTS exactly as it is (on-device wake < 300ms,弱网可用),只把"对话大脑"换成自有大模型,Overall response latency still stays within 1.5 seconds。

3. A four-step practical guide

Step 1: Align on requirements (1-2 weeks)

  • The brand's IT team aligns with the Nablai technical team on: how their own LLM will be deployed (private server / private cloud / VPC)
  • Character settings: the IP character's tone, values, conversation boundaries and how it handles sensitive topics
  • End-to-end SLA: response latency, availability and what triggers content moderation

Step 2: Private deployment (2-4 weeks)

  • The brand provides the LLM inference service (OpenAI-compatible API / custom protocol)
  • The Nablai module firmware is updated over OTA to a specified version and the private brain mode is enabled
  • Configure in the AMS console: API endpoint, authentication method and degradation policy

Step 3: Joint debugging (2-3 weeks)

  • Long-tail scenario testing: weak networks, packet loss, high concurrency, malformed input
  • Content moderation triggers: boundaries around self-harm, violence and politically sensitive topics
  • Character consistency testing: aligning tone and values across 100 representative conversations

Step 4: Staged rollout (4-8 weeks)

  • 10% of devices: monitor response latency, error rate and user activity
  • 50% of devices: monitor subscription conversion and the effect of content updates
  • 100% rollout: full OTA firmware push, with a seamless switch for users
Typical project timeline:From requirements alignment to 100% rollout: 2-4 months. It depends on how ready the brand's IT team is and how mature their own LLM is.

4. A typical architecture

[Toy-side LX002S module] On-device ASR (wake < 300ms)
   端侧 TTS(合成)
         ↓
   HTTPS / WebSocket
         ↓
[品牌方私有云 / VPC]
   自有大模型推理服务(OpenAI 兼容 API)
   内容审核模块
   角色库 / 故事库
         ↓
[梯度算子 AMS 后台]
   设备激活 / OTA / 用户活跃度
   订阅与计费(品牌方独立账户)

5. Four common questions

Q1: Is a private LLM expensive?

If you deploy an open-source model (Qwen, DeepSeek, Llama and others) privately, inference costs about ¥0.0001-0.001 per call at up to 100,000 QPS, one to two orders of magnitude cheaper than a public API. The one-time investment is GPU servers, roughly ¥300,000-1,000,000.

Q2: Will response latency get worse?

If your own LLM runs in a VPC in the same region, end-to-end latency stays within 1.5 seconds. Cross-region, it can reach 2-3 seconds. We recommend same-region deployment plus inference optimization (quantization, KV cache, prefix cache).

Q3: How is firmware OTA managed?

The Nablai AMS console supports pushing OTA independently for each brand. The module itself has staged rollout capability, so updates can be targeted precisely by device serial number, S/N range or user tag. A failed OTA rolls back automatically.

Q4: Who owns the subscription revenue?

The brand. Nablai takes no share of subscription revenue and charges only for the module BOM and an operations and maintenance fee. The complete brand-owned account system, subscriptions and content operations all sit with the brand.

6. When an LLM of your own is the wrong call

  • You only want one or two AI SKUs to test the market - running an MVP on a public API is faster
  • The brand's IT team cannot operate an LLM - strengthen the team first, or the private deployment becomes long-term debt
  • The target user base is small (< 5 万)——私有化投入回收周期太长

If what you are aiming for is long-term IP assets plus recurring subscriptions plus brand sovereignty, the LX Series with your own LLM is the safest combination in 2026. Full analysis: https://www.nablai.com.cn/articles/lx-series-brand-sovereignty.html

FAQ: Six questions you may have

1. Can an AI module connect to the brand's own LLM?
Yes. The Nablai LX Series lets you plug in the brand's own LLM (private deployment or private cloud) and ships with a complete brand-owned account system. On-device ASR/TTS is still handled locally by the LX Series module to keep wake-up fast, and once the conversation brain is swapped for your own LLM, end-to-end response latency still stays within 1.5 seconds.
2. Is connecting an AI module to your own LLM expensive?
If you deploy an open-source LLM (Qwen, DeepSeek, Llama and others) privately, a single inference costs about ¥0.0001-0.001 at up to 100,000 QPS, one to two orders of magnitude cheaper than a public API. The main one-time investment is GPU servers, roughly ¥300,000-1,000,000.
3. How long does a project to connect an AI module to your own LLM take?
A typical cycle is 2-4 months: 1-2 weeks for requirements alignment, 2-4 weeks for private deployment, 2-3 weeks for joint debugging and 4-8 weeks for staged rollout. It depends on how ready the brand's IT team is and how mature their own LLM is.
4. Who owns the user data once the AI module is connected to the brand's own LLM?
The brand. Nablai takes no share of subscription revenue and does not retain user conversation data. The complete brand-owned account system, subscriptions and content operations all sit with the brand.
5. Which open-source LLMs can be privately deployed with an AI module?
Anything that exposes an OpenAI-compatible API can be connected, including major open-source models such as Qwen, DeepSeek, Llama, ChatGLM and Baichuan, as well as public APIs such as Doubao, Tongyi, ERNIE, Kimi, GPT, Claude and Gemini (as a fallback).
6. What is the difference between the Nablai LX Series and TY Series when it comes to a proprietary LLM?
The LX Series is positioned for deep customization and brand sovereignty, with full support for private deployment of your own LLM. The TY Series is built on the Tuya platform with 60+ languages out of the box and is aimed mainly at global expansion; it connects to the Tuya cloud LLM, which already covers multiple languages, and does not support a brand's private LLM.

Shenzhen Nablai Intelligent Technology Co., Ltd. — AI module specialists for smart toys

📧 contact@nablai.com.cn    🌐 www.nablai.com.cn