China Merchants Bank Wins AI Efficiency Contest

News related to:China Merchants Bank · 2 min read
China Merchants Bank has emerged as a leader in cloud-native technology after winning the CNCF End User Case Study Contest for its innovative approach to unifying AI training and inference on Kubernetes. The bank's efforts have significantly improved the efficiency of its AI infrastructure, leading to a substantial increase in the utilization of its accelerator compute resources.
According to Chris Aniszczyk, CTO of the Cloud Native Computing Foundation (CNCF), China Merchants Bank's achievement highlights the potential of composable, vendor-neutral infrastructure to deliver measurable efficiency at scale. "As AI workloads continue to expand into critical infrastructure sectors such as finance, the core technology supporting it has to work harder without multiplying costs," said Aniszczyk. "China Merchants Bank's approach to unifying AI training and inference is a clear example of how this can be achieved."
China Merchants Bank built a unified control plane by integrating Kubernetes with a suite of CNCF projects, including Kueue, KEDA, Prometheus, HAMi, and Fluid. This framework allowed the bank's model training, fine-tuning, and online inference to share nearly 10,000 heterogeneous accelerator cards. As a result, the bank was able to increase the average utilization across these cards from 35% to more than 60%. Furthermore, the cost of processing 1 million tokens (input and output combined) was cut by more than 60% under comparable model and service conditions.
The bank's in-house Twinkle training framework also contributed to these improvements. By default, the framework allows five LoRA tenants to share one base model instance, reducing accelerator resource usage by 80% while increasing training density fivefold. This approach not only enhances the efficiency of the bank's AI infrastructure but also sets a precedent for other financial institutions looking to optimize their AI workloads.
"Built on a unified cloud-native foundation, our platform brings together heterogeneous compute pooling, training job scheduling, elastic inference scaling, accelerated access to data and models, and end-to-end observability," said PeiXiang Tan, AI Infrastructure Architect at China Merchants Bank. "This allows training and inference to follow workload-specific paths while working in close coordination, helping every unit of compute deliver greater value over time."
China Merchants Bank's achievement was recognized on stage during a keynote at KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China, further cementing its position as a trailblazer in the use of cloud-native technologies for AI workloads. The bank plans to continue building on these advances, focusing on key areas such as adjusting multi-tenant training concurrency dynamically, combining utilization, queue state, and latency signals into unit-cost-based capacity management, and extending KEDA toward serverless inference that can scale to zero between demand spikes.
By demonstrating the potential of unified AI training and inference on Kubernetes, China Merchants Bank has set a new standard for financial institutions looking to optimize their AI infrastructure. The bank's success highlights the importance of leveraging cloud-native technologies to enhance efficiency and reduce costs in the rapidly evolving field of artificial intelligence.