As global AI companies expand their operations in Singapore and enterprises worldwide invest in AI to drive business outcomes, demand for production-grade inference infrastructure is accelerating. To address this, Aolani, a Singapore-founded neocloud, has announced the launch of the Aolani Token Factory, a managed inference platform that enables organizations to deploy and scale AI models on a pay-per-token basis without the need to provision or manage underlying GPU infrastructure.
The launch positions Aolani as the first Singapore-founded neocloud to offer production-grade managed inference at scale. The platform is designed to close the accessibility gap, providing AI-native companies and enterprises a compliant and high-performance path from AI experimentation to production deployment.
The Aolani Token Factory operates on a per-token metering model, allowing customers to pre-purchase credits and pay based on token consumption, rather than investing in capital-intensive GPU infrastructure. Aolani manages the entire inference stack, including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization. This enables customers to scale consumption without continuously provisioning additional infrastructure.
At launch, the platform supports leading open-source models including DeepSeek, GLM, Kimi, and Qwen. Aolani plans to expand the model catalog over time based on customer demand. Customers can also deploy their own models through OpenAI-compatible APIs. For enterprises with strict compliance and data residency requirements, dedicated capacity and data isolation options are available.
The Aolani Token Factory supports three core production use cases: AI agents for high-volume inference and workflow automation; enterprise AI applications such as internal copilots, knowledge assistants, and document intelligence; and coding agents for code generation, completion, testing, and review.
Sea Xu, Applied AI Research Lead at Aolani, said: "The Aolani Token Factory is built on a high-performance inference stack that supports the most in-demand open-source model families. We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge. As Southeast Asia's AI ecosystem evolves and grows rapidly, it is our goal to ensure that the infrastructure serving it keeps pace."
Nicholas Chia, Chief Executive Officer at Aolani, added: "Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure. This is a significant milestone for us and our customers as the Aolani Token Factory will fundamentally change how customers access AI compute."
The launch comes at a time when companies across Southeast Asia are seeking to leverage AI to gain competitive advantage. By lowering the barrier to entry for AI inference, the Aolani Token Factory enables organizations of all sizes to incorporate AI into their operations without significant upfront investment. This move is expected to accelerate AI adoption in the region and support the growth of the AI ecosystem in Singapore.
For more information about the Aolani Token Factory, visit Aolani's Token Factory page. You can also explore Aolani's broader offerings at their website or follow them on LinkedIn.

