The market gap hiding in plain sight
The global AI race has two lanes: English and Mandarin. Every major frontier lab — OpenAI, Anthropic, DeepSeek, Moonshot AI, Z.ai — is headquartered in either the United States or mainland China, and their models reflect that geography. For the more than 80 million people who speak Cantonese, roughly the same number as Korean speakers and more than Italian or Thai speakers combined, that dominance has a practical cost.
'The whole AI revolution is in English and Mandarin,' says Pak-Sun Ting, CEO of Votee AI, a Hong Kong-based startup building language models for Cantonese and other underserved tongues. 'There's only a very small fraction that represents other languages.'
What Votee actually does
Votee does not build from zero. It takes open-weight models from developers like Meta and Alibaba, retrains them on Cantonese data, and sells the finished product to banks, universities, and government departments. The Cantonese corpus Votee has assembled grew from 100 million tokens to more than 500 million, drawn from online scraping — including content from Radio Television Hong Kong — university partnerships, community contributions, and synthetic data sets the company generates itself.
The resulting models run at around 70 billion parameters, significantly smaller than frontier offerings. Training costs land at roughly $250,000 per run, using between 500 million and 1 billion tokens, compared to the trillions consumed by English-language frontier models. The numbers come first: this is a lean operation competing on precision, not scale.
Ting is direct about the stakes. 'It's more than just culture. Cantonese is used in education, healthcare, and police communications,' he says. 'If those don't get covered, then AI is essentially useless.'
HKCanto-Eval, a benchmark set developed by researchers at Kyushu University, the Education University of Hong Kong, and the local AI community hon9kon9ize — and sponsored by Votee — confirms the gap. Mainstream models handle everyday Cantonese at a reasonable level but routinely fail on cultural and local knowledge.
Sovereign AI: the bigger frame
Votee is not alone. Indonesia's Indosat is building Sahabat AI for Bahasa and related languages. Singapore's state-backed AI Singapore runs SEA-LION across 11 Southeast Asian languages. South Korea has staged a state-sponsored model competition — local media dubbed it the 'AI Squid Game' — backed by a 2026 AI budget of roughly $6.8 billion.
All of it travels under the label 'sovereign AI': the principle that governments and enterprises should own their data, models, and infrastructure rather than rent capability from a foreign provider who can, as Ting puts it, 'turn it off.'
Ting is candid that full supply-chain sovereignty is 'very difficult.' His practical prescription: own the foundation models and the applications built on top of them. For routine government automation, even a model as small as 1 billion parameters can be sufficient.
The editorial read
The sovereign AI argument is, at its core, a free-enterprise and national-security argument dressed in technical language. Governments and private institutions that outsource their core AI infrastructure to a handful of U.S. or Chinese hyperscalers are trading long-term autonomy for short-term convenience — a bargain that looks worse the more geopolitically turbulent the environment becomes.
Votee's model is instructive precisely because it is capital-efficient. A $250,000 training run is not a rounding error for a startup, but it is a number that governments, universities, and regional banks can actually budget for. The market has already voted: when the cost of sovereignty drops low enough, buyers show up. That is how durable alternatives to concentrated power get built.



