# KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China — Schedule

Machine-readable mirror of the schedule page. Generated 2026-09-16.

- Source: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/
- Sessions: 129 across 3 days
- Times are the event's local times, exactly as published. Timing and rooms are subject to change.
- Each session links back to the schedule page, which opens that session's details.

## Tracks

- Keynote Sessions (24)
- Project Opportunities (15)
- AI + ML + Agentic AI + Data Systems (10)
- AI Infrastructure + Accelerators + Performance Engineering (9)
- Maintainer Track (8)
- ⚡Lightning Talks (8)
- Registration + Badge Pick-up (6)
- Breaks (6)
- Cloud Infrastructure + Virtualization + Storage (6)
- Experiences (5)
- Emerging Technologies + Research + Advanced Topics (4)
- Sponsor-hosted Co-located Events (3)
- Sponsor Demos (3)
- Community + Open Source + Getting Started (3)
- Networking + Edge + Distributed Systems (3)
- Application Development + Developer Experience (3)
- Operations + Observability + Reliability (3)
- Solutions Showcase (2)
- Platform Engineering + Cloud Native Architecture (2)
- 📚 Tutorials (2)
- Security + Privacy + Trusted Computing (2)
- CNCF-Hosted Co-located Events (1)
- Hardware Enablement + Diversification (1)

## Monday, September 7, 2026

### 08:00–17:15 · Registration + Badge Pick-Up

- Room: 1F Foyer
- Track: Registration + Badge Pick-up
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1283926

### 08:00–17:15 · Cloakroom

- Room: 7F Foyer
- Track: Registration + Badge Pick-up
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1283930

### 13:00–17:00 · Open Computing: Building Open Software for Diverse Hardware Hosted by FlagOS

- Room: 5F | 5B + C
- Track: Sponsor-hosted Co-located Events
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1312720
- Website: http://bagevent.com/event/flagos

From September 7–9, 2026, KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China will be co-located in Shanghai, expected to draw over 1,000 developers, architects, and technical decision-makers, alongside 400+ participating companies. As a flagship co-located event, Open AI Computing: Building an Open-Source Software Ecosystem for Diverse Heterogeneous Hardware will be held on September 7. It's a high-caliber platform for AI framework developers, compiler and runtime researchers, open-source infrastructure builders, and hardware architects to exchange ideas.

AI chip diversification is now an industry reality, yet a fragmented software stack ecosystem has become the critical bottleneck preventing compute from scaling into large-model and agent applications. Every new chip demands repeated porting of operators, compilers, and inference engines; every framework must independently contend with the long tail of multiple backends. In this multi-chip era, a unified AI system software stack is essential. FlagOS is an open-source AI system software stack for multi-chip scenarios, spanning a unified operator library, a multi-backend AI compiler, training and inference frameworks, a communication library and other core components. As the host of this event, the FlagOS community will bring together the most representative experts from Beijing Academy of Artificial Intelligence (BAAI), Shanghai AI Laboratory, PyTorch Foundation, vLLM, SGLang, NVIDIA,etc. to jointly explore the technical roadmaps and collaboration models of an open AI system software stack—showcasing new system software technologies that span multiple chip architectures, and discussing hardware–software co-optimization approaches across emerging AI chip architectures, open-source compilers and acceleration frameworks.

For questions regarding the event, please contact: contact@flagos.io

Registration will open by August 26, 2026.

### 13:30–18:00 · OSPOlogy + OSPO Summit China 2026 | [Pre-Registration Required]

- Room: 5F | Jiu Zhou Hall
- Track: CNCF-Hosted Co-located Events
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1308451
- Website: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/co-located-events/ospology-ospo-summit/

The OSPOlogy + OSPO Summit China 2026 focuses on the evolution of Open Source Program Management (OSPO) functions and corporate open-source governance in the era of Agentic AI. The conference will explore key topics including how Agentic AI supports practical OSPO activities, Large Language Model (LLM) and data governance, AI-driven software supply chains and enterprise global expansion strategies as well as open strategies and innovation mechanisms.

Don't miss out on this CNCF-hosted co-located event, register today! https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/co-located-events/ospology-ospo-summit/#registration-details

### 13:30–18:30 · AI Native Open Session Hosted by ETSI ISG NFV

- Room: Off-Site Location
- Track: Sponsor-hosted Co-located Events
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1312708
- Website: https://nfvwiki.etsi.org/index.php?title=2nd_AI_Native_Open_Session_Hosted_by_ETSI_ISG_NFV_%E2%80%93_7_Sep_2026_%E2%80%93_Shanghai

Please note that this is an off-site Sponsor-hosted Co-located event.

Location: Pudong Shangri-La, Shanghai

CNCF has led the revolution of cloud native for the IT industry, and telecommunications have followed this direction since CNCF issued the definition of cloud native. Organizations like ETSI NFV ISG have aligned with CNCF’s flagship direction, and global telecom operators have deployed their 5G networks with a cloud‑native vision.

With the rapid growth of AI and agent technologies, cloud‑native platforms offer a scalable, reliable, and standardized foundation for modern applications. CNCF is proactively defining and leading the AI‑Native paradigm through open standards and projects, ensuring future software development efficiently leverages artificial intelligence at scale.

ETSI NFV ISG held the first AI‑Native Open Session at KubeCon + CloudNativeCon Europe 2026 and encouraged CNCF to continue leading this direction. On 7 September 2026, during KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch + AgntCon China 2026, ETSI NFV ISG will organize the 2nd AI Native Open Session, gathering global AI‑native leaders in Shanghai.

ETSI NFV ISG held the first AI native summit and encouraged CNCF to continue to lead ‘AI Native’ direction at Kubecon EU Amsterdam 2026.

Attendance is free of charge.

For questions regarding the event, please contact: denghui02@hotmail.com

### 14:00–16:00 · Hardware-Aware AI: Building PyTorch Workloads Across Accelerators Hosted by Huawei

- Room: 7F | Pearl Hall
- Track: Sponsor-hosted Co-located Events
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1312719

Driven by the vigorous growth of diverse computing resources, AI workloads are required to run seamlessly on a wide range of heterogeneous accelerators such as GPUs, NPUs and XPUs. Centered on hardware‑aware AI, this forum focuses on building the cross‑accelerator ecosystem for PyTorch. It shares hands‑on industry practices covering framework‑level adaptation, operator implementation, compiler co‑optimization, workload migration and performance tuning, and jointly explores the practical challenges and future technical pathways for integrating domestic heterogeneous computing into the upstream PyTorch community.

For questions regarding the event, please contact: zesheng.zong@outlook.com

Registration will open by August 26, 2026

### 17:30–20:30 · PyTorch Conference Community Night, Sponsored by Huawei

- Room: Off-Site Location
- Track: Experiences
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1314315

Location: Puzhixing, F4, Pujiang Hall
(No. 1 Dongchang Road, Pudong New Area, Shanghai)

Join fellow developers, researchers, and AI builders for a fun evening of great conversation, food, and drinks. It’s the ultimate way to gear up with the community before the main conference begins tomorrow.

Space is limited, available on a first come, first entered basis. KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China badge required for entry and must be picked up from the conference venue, Shanghai International Convention Center.

RSVP: https://linuxfoundation.research.net/r/PyTorchReception

## Tuesday, September 8, 2026

### 07:30–18:30 · Registration + Badge Pick-Up

- Room: 1F Foyer
- Track: Registration + Badge Pick-up
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255197

### 07:30–19:30 · Cloakroom

- Room: 7F Foyer
- Track: Registration + Badge Pick-up
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255200

### 09:00–09:02 · From Open Source Adoption to Open Source Innovation: China’s Next Chapter

- Room: 7F | Grand Ballroom II + III
- Speakers: Professor Lu Shouqun
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1308499

Drawing on decades of experience advancing open source in China, Professor Lu Shouqun reflects on the evolution of China’s open source ecosystem and its growing connection to the global community. He explores how open source has progressed from technology adoption to community-driven innovation and why collaboration across borders, projects, and emerging technologies will be increasingly important as AI reshapes the industry.

### 09:02–09:07 · Open Source for the AI Era

- Room: 7F | Grand Ballroom II + III
- Speakers: Jonathan Bryce, Horace Li
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1304838

This opening will introduce the complete open source AI stack, from models and frameworks to inference, cloud native operations, infrastructure and hardware, and set the stage for the stories that follow. Through real-world examples across both days, the audience will see how these technologies work together to build, run and scale frontier intelligence—and how open collaboration across the stack is shaping what comes next.

### 09:07–09:10 · Beyond Multimodal: Building Full‑Modal AGI via Open‑Weight

- Room: 7F | Grand Ballroom II + III
- Speakers: Ryan Lee
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1313048

### 09:12–09:22 · Operating Frontier Intelligence at Scale

- Room: 7F | Grand Ballroom II + III
- Speakers: Chris Aniszczyk, Xiao Zhang
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1312064

"Building a model is only the beginning. When AI moves into production, the challenge shifts to making every GPU count, scaling dynamically with demand, maintaining reliability, and understanding increasingly complex systems. Cloud native technologies are becoming the operating layer that makes this possible.

We’ll also look at the growing role of observability as AI systems become more complex and why the next generation of AI infrastructure will depend on open technologies working together across the stack."

### 09:24–09:30 · Tidal Auto-scaling for Training and Inference Based on Kubernetes + KEDA

- Room: 7F | Grand Ballroom II + III
- Speakers: Jun Zheng, Tan Pei Xiang
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1317065

CMB AI assistants exhibit distinct peak and off-peak traffic patterns. During off-peak hours, the platform allocates idle GPUs to LoRA incremental training. As peak traffic arrives, KEDA scales out vLLM inference services based on request load, prompting training tasks to save progress at secure checkpoints and proactively release capacity. Once traffic subsides, training automatically resumes from the saved checkpoints. This solution prioritizes online services to meet latency targets while continuously leveraging off-peak resources for model iteration, achieving coordinated optimization between inference SLOs and training efficiency.

### 09:32–09:42 · Building Frontier Intelligence in the Open

- Room: 7F | Grand Ballroom II + III
- Speakers: Mark Collier
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1312046

The pace of AI innovation requires an open ecosystem where researchers, developers, model builders and hardware companies can innovate together and where breakthroughs can move quickly from experimentation to production.

This keynote explores how PyTorch has become a foundation for that ecosystem, enabling intelligence to be built and optimized across a rapidly expanding landscape of models and compute. From open models to new accelerator architectures, we'll look at how an open development ecosystem expands who can build AI, how quickly they can innovate, and how that intelligence connects to the infrastructure that ultimately runs it. Huawei sponsored keynote to be included We’ll also welcome the PyTorch Foundation’s newest Platinum Member, marking a significant new investment in the open ecosystem and the future of AI hardware innovation.

### 09:42–09:47 · Ascend & PyTorch: Pioneer New AI Open Ecosystem

- Room: 7F | Grand Ballroom II + III
- Speakers: Liang Zhang
- Track: Keynote Sessions
- Labels: Chinese, Sponsored
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1317073

As AI models grow larger and their architectures become increasingly sophisticated, the industrial ecosystem faces new opportunities and challenges. Drawing on eight years of opensource practice, Ascend deepens its collaborative partnership with the PyTorch community, evolving from hardware adaptation to joint research on technical standards.

Together, we explore nextgeneration AI hardwaresoftware codesign solutions, foster interoperability between Chinese and global AI technology stacks, and build an open and sustainable AI industrial ecosystem.

### 09:47–09:52 · Towards Device-agnostic PyTorch: Building Unified Infrastructure for a Multi-Backend Ecosystem.

- Room: 7F | Grand Ballroom II + III
- Speakers: Wei Li
- Track: Keynote Sessions
- Labels: English, Sponsored
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1313965

AI computing has moved from chips to platforms, and now to ecosystems. Cambricon's decade-long journey has proven that scale demands full-stack co-optimization—and we are taking that expertise upstream. With an upstream-first philosophy and a track record of real contributions, we are hardening PyTorch's device-agnostic foundation for every PrivateUse1 backend. Our goal: fixes upstream, seamless everywhere. Device-agnostic is just the enabler; the real prize is unlocking what each backend does best—broader device reach and richer backend-native capabilities.

### 09:54–09:57 · Inside vLLM: Production Best Practices, Model Integration and Road Map

- Room: 7F | Grand Ballroom II + III
- Speakers: Kaichao You
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311164

Take a deep dive into the vLLM ecosystem to see how to optimize, scale, and extend your LLM serving infrastructure. Beyond essential production deployment strategies, we will unpack major platform updates, including the Rust frontend, KV connector, and reinforcement learning rollouts. Finally, learn the practical mechanics of integrating brand-new architectures, with Kimi K3 as an example.

### 09:59–10:04 · PD Disaggregation vLLM Deployment on Alternative AI Accelerators Using llm-d

- Room: 7F | Grand Ballroom II + III
- Speakers: 纪飞 王, Mengxuan Li
- Track: Keynote Sessions
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1313959

### 10:06–10:11 · From Containers to Agents: The Next Cloud Native

- Room: 7F | Grand Ballroom II + III
- Speakers: Di Xu
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1309269
- Slides: https://sessionize.com/download/iicosna~HupukHSbEyx51Co8oHcD6Y.pdf~kubeconch26-keynote-from-containers-to-agents-the-next-cloud-native-abstraction.pdf

Every major computing shift creates a new unit of work.
Virtual machines gave us infrastructure automation. Containers gave us portable applications. Kubernetes gave us a control plane for distributed systems.
Agents may become the next unit of work.

An agent is not just an LLM call. It has identity, memory, goals, tools, permissions, execution state, and the ability to affect the outside world. It can fail in ways that traditional software does not: by taking the wrong action, using the wrong tool, leaking context, exceeding its authority, or behaving differently as its models and environment evolve.

The industry is rapidly building agent demos. The harder problem is operating agents in production.

This keynote proposes a cloud native view of agent infrastructure: agents as first-class workloads with declarative configuration, policy-driven tool access, observable decision paths, isolated execution, durable state, continuous evaluation, and human-in-the-loop controls.

Rather than treating agent infrastructure as a proprietary layer above Kubernetes, we have an opportunity to apply the lessons of cloud native computing—open interfaces, reconciliation loops, composable primitives, and policy as code—to make autonomous systems safer and more portable.

The question is no longer whether agents will run in production. The question is whether we will build the infrastructure needed to trust them there.

### 10:13–10:18 · What AI Agents Need from Open Infrastructure

- Room: 7F | Grand Ballroom II + III
- Speakers: Yaya Xia
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1308539

AI agents create a demanding kind of workload. They generate and run code, interact with external systems, and often need an isolated execution environment for only a few minutes.
The cloud native and open infrastructure communities have spent years building the foundations this workload now needs. This keynote examines how those existing capabilities can be assembled into an agent runtime, supported by new ecosystem evidence from CNCF and OpenInfra.

The idea becomes concrete through an open-source runtime stack. Kubernetes Agent Sandbox manages the sandbox lifecycle, while Kata Containers provides workload isolation, with other cloud native components supporting execution and resource delivery. Together, these existing technologies form an agent runtime designed for the needs of agentic AI.

### 10:20–10:25 · Building an Agent Runtime with Open Infrastructure

- Room: 7F | Grand Ballroom II + III
- Speakers: Yaya Xia, Xu Wang
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1309274

An agent runtime has to create a working environment on demand, isolate code the platform may not trust, and release resources when the task ends. Much of this machinery already exists across cloud native and open infrastructure projects.

This keynote follows one agent task through an open-source runtime stack.
Kubernetes Agent Sandbox manages its lifecycle, while Kata Containers provides a dedicated guest-kernel boundary. The rest of the chain handles container execution and efficient delivery of images and model artifacts. Seen together, these projects offer a practical foundation for secure, short-lived agent workloads and a shared path for the ecosystem to improve it.

### 10:25–10:30 · Closing Remarks

- Room: 7F | Grand Ballroom II + III
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1304846

### 10:30–11:00 · Coffee Break ☕

- Room: 7F | Grand Ballroom I
- Track: Breaks
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255188

### 10:30–11:00 · Gold Sponsor In-Booth Demos

- Room: 7F | Grand Ballroom I
- Track: Sponsor Demos
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1319134

Sponsor: LF Open Source Software Academy (LFOSSA)
Demo: Master AI & Open Source Skills and to Advance Your Career in the Age of AI
Booth Number: G5

In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.

### 10:30–14:30 · Project Pavilion | Tuesday AM Project Tables

- Room: 7F | Grand Ballroom I
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311154

Tuesday AM Project Tables | 10:30 - 14:30

T-1 HAMi
T-2 Karmada
T-3 Argo
T-4 Dragonfly
T-5 CNCF Table
T-6 Community-Led Events
T-7 Cilium
T-8 Kmesh
T-9 Fluid
T-10 Atlantis
T-11 StarlingX (OpenInfra project)
T-11 Zuul (OpenInfra project)
T-11 OpenStack (OpenInfra project)
T-12 PyTorch
T-13 LWS

View the full project table directory here: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/features-add-ons/project-engagement/#project-table-directory

### 10:30–19:00 · Solutions Showcase

- Room: 7F | Grand Ballroom I
- Track: Solutions Showcase
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255194

### 10:30–11:00 · Women's Community Gathering

- Room: 7F | Grand Ballroom I
- Track: Experiences
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311104

Grab coffee, tea, and a snack during the morning break and meet other women, non-binary, and allies to network, connect and socialize. Look for the Reserved Sign on a table in the Solutions Showcase.

Women’s Gatherings enable meaningful networking, peer mentorship, collaboration, and sustained engagement.

### 11:00–11:30 · 8 Million Tasks Daily: Driving Efficiency in Multi-Tenant Data Platforms at Horizon Robotics

- Room: 1F | Mandarin Hall I
- Speakers: JIanxiang Sun, Shuangkun Tian
- Track: Platform Engineering + Cloud Native Architecture
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1225136
- Slides: https://sessionize.com/download/ccagmeh~Sho5sytNjC4agwxxVJ33VD.pdf~kubecon_shanghai_2026_horizon_robotics.pdf

As autonomous driving shifts to end-to-end models, Horizon Robotics faces unprecedented scale. Thousands of concurrent, sub-minute workflows pushed infrastructure to its limits, where traditional orchestration and "Informer Choke" caused severe instability and resource waste.
We re-architected our platform to handle 8 million daily tasks, saving 1,000+ CPU nodes. We will dive into:
1. Orchestration Re-engineering: Replacing per-workflow agents with a centralized Argo Workflows Agent Service. This "Agentless" approach cut infrastructure costs by 2.4M CNY/month and saved 1,000+ CPU nodes.
2. Scheduling & State Optimization: Overcoming Informer bottlenecks by refactoring metadata, reducing VolcanoJob size by 30%, and slashing sync latency from 4 minutes to <10 seconds.
3. Stability & Multi-Tenancy: Leveraging Kueue and Volcano to implement elastic quotas and resolve preemption conflicts, ensuring system stability in large-scale multi-tenant environments.

### 11:00–11:30 · Accelerating Sequential Recommender Models in the PyTorch Ecosystem

- Room: 1F | Mandarin Hall II
- Speakers: Zan Huang, Jiayu Sun
- Track: Community + Open Source + Getting Started
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1210254
- Slides: https://sessionize.com/download/izuxis~nWSBcPQMv3MwVn8Qb4baGS.pdf~pytorchchina2026seqrecopt.pdf
- GitHub Repository: https://github.com/pmixer/SASRec.pytorch/
- Project Demo: https://github.com/NVIDIA/cudnn-frontend/tree/develop/python/cudnn/hstu_attention
- Website: https://github.com/pytorch/FBGEMM/commit/e0dec03942b7373d9d714c70bfc77e3b44a694e8

Sequential recommendation models share a natural connection with large language models in use of self-attention, yet production recommendation systems introduce meaningful differences that are often underappreciated.

We begin with SASRec, one of the earliest Transformer decoder-based sequential recommenders, walking through its rewrite from TensorFlow to PyTorch - highlighting how PyTorch reduced modeling complexity and surfacing the gaps between academic prototypes and production systems.

We then examine HSTU, an attention operator for sequential recommendation at scale: what makes its attention pattern structurally different from standard Transformer attention, and key GPU insights - kernel fusion, variable-length sequences, and PyTorch integration via FBGEMM.

We close with a reflection on sequential recommendation and LLMs from modeling and systems perspectives, sharing observations on emerging trends and optimization strategies across recommendation workloads.

### 11:00–11:30 · Declarative Underlays: Scaling Purpose-Built Infrastructure(Clusters) for OpenStack with Cluster API

- Room: 7F | Grand Ballroom II + III
- Speakers: HanSol Park, Kangsub Song
- Track: Cloud Infrastructure + Virtualization + Storage
- Labels: Advanced, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224992
- Slides: https://sessionize.com/download/iktariq~RSKSn2UTCwcWzwsfmdNFr2.pdf~declarative_underlays_kubecon_china_2026.pdf

OpenStack and k8s were once considered separate ecosystems, but today many operators run OpenStack on top of Kubernetes to get the benefits of both worlds.

We have taken that integration a step further by leveraging Cluster API (CAPI)
to declaratively create/manage clusters, each tuned for a specific need (hardware/software)
(AI clusters /w GPUs, high-speed networking cluster, dedicated storage, etc.).

The underlay becomes a set of clusters that can be re-balanced on demand,
delivering true cloud-native agility to the physical layer.

In this session, we will focus on production-proven architecture that:
- Declarative underlay clusters – repeatable, version-controlled provisioning.
- Isolates specialized hardware – Cluster for (AI, Compute, Network) without interfering with other workloads
- Dynamically rebalances resources – migrate idle nodes (GPUs) between clusters
- Integrates OpenStack services
- Operates at scale (Multi Region) – lessons learned and best-practice patterns

### 11:00–12:15 · 📚 Building Portable Large Models on Heterogeneous AI Accelerators with FlagOS

- Room: 7F | Pearl Hall
- Speakers: Mengsi Lyu, Hongjun Zhang, Yufeng Lyu
- Track: 📚 Tutorials
- Labels: AI Infrastructure + Accelerators + Performance Engineering, Any, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224095

Porting large models across diverse accelerators and running them efficiently remains difficult. Developers can download LLMs and multimodal models such as Qwen, MiniMax, and DeepSeek, yet the runtime path below PyTorch, vLLM, or SGLang is still fragmented across vendor libraries, compiler stacks, communication backends, and framework patches.
This tutorial introduces how to port large models onto various AI accelerators by using the FlagOS open-source stack. we will walk through the path from PyTorch operators to training and inference with vLLM and SGLang.
Audiences will learn how Triton enables portable kernel programming, how FlagGems provide PyTorch operator coverage, how unified plugins of FlagOS for vLLM, Megatron, VeRL etc. support training and inference on different AI accelerators, and KernelGen help auto-coding for different accelerators. This tutorial shows how developers can bring open models to more hardware, with less porting effort and stronger community collaboration.

### 11:00–11:05 · Project Lightning Talk: Opening Remarks

- Room: 5F | 5B + C
- Speakers: Miley Fu
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1260073

### 11:07–11:12 · Project Lightning Talk: Architecture Without Gatekeepers: Decoupling for Neutrality

- Room: 5F | 5B + C
- Speakers: Sergey Pronin
- Track: Project Opportunities
- Labels: English, OpenEverest
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1223701

Open source only works if the community can contribute without hitting a wall. We realized our architecture at OpenEverest was unintentionally biased toward specific tools, so we rebuilt it. I’ll share how we moved to a modular plugin system that decouples our core from specific operators. By supporting everything from Valkey to competing Postgres flavors, we’ve eliminated vendor lock-in and ensured the community—not the architecture—dictates the project's future.

### 11:14–11:19 · Project Lightning Talk: From Static Slices to Elastic GPUs: Dynamic MIG with HAMI

- Room: 5F | 5B + C
- Speakers: 纪飞 王
- Track: Project Opportunities
- Labels: Chinese, HAMi
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1215494

In Kubernetes, NVIDIA MIG is typically pre-partitioned, forcing operators to decide GPU slice layouts before workloads arrive. This static model struggles with mixed, dynamic workloads, often leading to fragmentation and underutilization.

This talk presents a scheduling-driven approach: instead of fitting workloads into fixed slices, HAMI adapts GPU partitioning based on real-time scheduling decisions by integrating the scheduler with the device plugin.

The result is a more flexible and efficient use of GPU resources—while still operating within hardware constraints.

The key idea is simple: GPU partitioning should follow scheduling, not precede it.

### 11:21–11:26 · Project Lightning Talk: KubeEdge Everywhere: Latest Project Update with industrial cases

- Room: 5F | 5B + C
- Speakers: Hongbing Zhang
- Track: Project Opportunities
- Labels: English, KubeEdge
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224863

Since becoming the first cloud-native edge project to graduate from the CNCF, KubeEdge adoption has exploded. From intelligent transportation and smart cities to energy and banking, it is now the backbone of critical edge-cloud ecosystems. In this lightning talk, we will highlight the latest post-graduation features, governance updates, and real-world success stories that demonstrate the power of KubeEdge for Edge AI and industrial workloads.

### 11:28–11:33 · Project Lightning Talk: Multi-cluster Orchestration System: Karmada Updates and Use Cases

- Room: 5F | 5B + C
- Speakers: Hongcai Ren, Yiheng Ci
- Track: Project Opportunities
- Labels: Chinese, Karmada
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222021

Karmada (Kubernetes Armada) is a Kubernetes management system that enables you to run your cloud-native applications across multiple Kubernetes clusters and clouds.

In this presentation, the maintainer of the Karmada project will share:

- A Brief introduction to Karmada.
- New features over the last year
- Real-world case studies
- Overview of the community
- Roadmap

### 11:35–11:40 · Project Lightning Talk: Atlantis: Terraform Pull Request Automation for Cloud Native Teams

- Room: 5F | 5B + C
- Speakers: Rui Chen
- Track: Project Opportunities
- Labels: Atlantis, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1218468

Atlantis is a CNCF Sandbox project that provides Terraform and OpenTofu pull request automation for cloud native teams. In this 5-minute Project Lightning Talk, I will give a concise project update and show how Atlantis fits into modern infrastructure delivery: making plans visible in pull requests, keeping applies approval-gated, and preserving an auditable Git-based workflow.

The talk will focus on the core Atlantis workflow: webhook-triggered plans, pull request comments, policy-aware approvals, controlled applies, and self-hosted deployment patterns for Kubernetes-based teams. I will also briefly cover why this workflow remains important as the open source IaC ecosystem evolves around Terraform and OpenTofu.

Attendees will leave with a practical understanding of what Atlantis does, when it is a good fit, recent community focus areas, and how to get involved as users or contributors.

### 11:42–11:47 · Project Lightning Talk: Run Sandbox Securely and Cost-Effectively with OpenKruise Agents

- Room: 5F | 5B + C
- Speakers: Zhang Zhen
- Track: Project Opportunities
- Labels: OpenKruise
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1319491

Running sandboxes for LLM-generated code in production forces an uncomfortable three-way trade-off between security, latency, and cost. The sandbox executes untrusted model code, so it cannot hold a single outbound credential; a warm pool delivers sub-second creation but burns money at rest; and the agents inside must answer user requests and scheduled jobs instantly, even though they sit idle most of the time. This talk shows how OpenKruise Agents squares that triangle by combining horizontal and vertical elasticity on the warm pool, on-the-fly token injection through an agent gateway, and smart sandbox pause-and-resume. The takeaway: with Envoy and OpenKruise in-place update, production-grade agent sandboxes are already within reach.

### 11:45–12:15 · Scaling Digital Employees at China Merchants Bank in Production: A Controlled AI Agents Approach

- Room: 1F | Mandarin Hall I
- Speakers: Jiahang Xu
- Track: Platform Engineering + Cloud Native Architecture
- Labels: Advanced, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1225364

China Merchants Bank has transitioned from AI prototypes to large-scale internal deployment of Large Language Models, AI Agents for Digital Employees. To manage the inherent unpredictability of these autonomous systems in the regulated banking sector, the bank utilizes an “In-Control” architectural pattern built on Platform Engineering principles for its internal operations. This framework rests on three pillars: Platform Engineering for AI: Streamlines internal deployment using standardized “AI Agent-as-a-Service” templates to empower Digital Employees. Secure Sandboxing: Employs Kubernetes-native micro-VMs to execute agent-generated code in isolation, protecting core infrastructure. Financial Guardrails: Implements real-time monitoring and deterministic policies to ensure compliance and budget control within the enterprise. This structured approach allows the bank to manage internal Digital Employees in cloud environments while maintaining strict data sovereignty and auditability.

### 11:45–12:15 · Dragonfly's CNCF Graduation Journey and Large-Scale AI Model Distribution in Practice

- Room: 1F | Mandarin Hall II
- Speakers: Fengjun Lyu, Wenbo Qi
- Track: Community + Open Source + Getting Started
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1189573

In 2025, after nearly eight years of incubation within CNCF, the Dragonfly project officially graduated. But getting there wasn't easy. Growing from an internal project at Alibaba to a full-community-driven CNCF graduated project came with its fair share of challenges. In this talk, we will discuss the inner workings of Dragonfly's journey to its current state.

We’ll cover our lessons in community governance, including how we evolved from a small group of maintainers to transparent decision-making and attracted global contributors through effective open-source practices.

On the technical side, the increasing popularity of LLMs has made the efficient distribution of model files at the GB or even TB scale a significant challenge. We'll dive into how Dragonfly's P2P capabilities have evolved to tackle the challenge of pulling massive model files concurrently across Kubernetes clusters, along with best practices we've developed for AI inference and training workloads.

### 11:45–12:15 · Exploring HiFloat8: A Tapered Format Complementing the FP8 Ecosystem for Robust Model Training

- Room: 7F | Grand Ballroom II + III
- Speakers: Yun Zhao
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1185605
- Slides: https://sessionize.com/download/irpairze~T3bdM6bbezC1rtgv8nPTWZ.pdf~pytorchcon-china.pdf

Standard FP8 formats suffer from frequent gradient overflows and heavy reliance on complex Delayed Scaling, which often lead to training instability or suboptimal convergence in large models. This session introduces HiFloat8 (HiF8) — a tapered precision format that offers an alternative approach to managing dynamic range. This "natural" alignment with neural network weight/gradient distributions allows HiF8 to capture high-magnitude outliers without the aggressive scaling required by standard FP8.We explore how HiF8 can works in the training and inference procedure.We will demonstrate the implementation of HiF8 within the ecosystem on vllm and
deepspeed. They allow developers to evaluate performance of Hif8 on GPUs. Also, we will give an analysis of training stability and final loss parity where HiF8 provides relatively the same accuracy and 1.5-1.7 times GEMM performance than FP16. Finally, we will share insights from our ongoing collaboration on dedicated hardware support for HiF8.

### 11:49–11:54 · Project Lightning Talk: Let's Talk About Perses, the Open Specification for Prometheus and More

- Room: 5F | 5B + C
- Speakers: Leon Nunes
- Track: Project Opportunities
- Labels: Perses
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1319502

In this talk we will discuss perses a CNCF sandbox project, which is an open specification for Dashboards, perses can help you organize your data-sources, has embeddable components is devops ready and Kubernetes native. I'll be sharing what the Project is and what the latest is on this project in this short talk along with a small demo.

### 11:53–11:58 · Project Lightning Talk: Let’s Talk About Headlamp, the Extensible Open-Source Kubernetes UI

- Room: 5F | 5B + C
- Speakers: Emily Chen
- Track: Project Opportunities
- Labels: Headlamp
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1322749

In this talk, we will explore Headlamp, a CNCF Sandbox project that provides a modern and user-friendly interface for Kubernetes. The session will cover what Headlamp is, its key features and extensibility, recent developments, and how it fits into the Kubernetes ecosystem. This short talk will provide an introduction to the project and what to know when getting started.

### 11:58–12:00 · Project Lightning Talk: Closing Remarks

- Room: 5F | 5B + C
- Speakers: Miley Fu
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1260075

### 12:15–13:45 · Lunch (Not Provided) 🍜

- Room: Off-Site Location
- Track: Breaks
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255193

### 13:45–14:15 · Architecting Secure Agentic Workflows on Kubernetes: A Financial Sector Case Study

- Room: 1F | Mandarin Hall I
- Speakers: Vincent Caldeira, Morgan Foster
- Track: AI + ML + Agentic AI + Data Systems
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1177779

As organizations move beyond basic LLM integrations toward autonomous agentic workflows, the infrastructure required to support these systems grows increasingly complex. Running multi-agent architectures in production introduces unique challenges around tool discovery, secure access, and traffic routing.

This session explores how to leverage Kubernetes and cloud-native abstractions to develop, test, and deploy AI agents at scale. Using a reference architecture leveraging the Kagenti project, we will demonstrate how to construct a secure, scalable agentic environment. The talk covers unifying tool access via an MCP Gateway, enforcing zero-trust workload identity with SPIFFE/SPIRE, and standardizing inter-agent communication.

To illustrate these concepts, we will walk through a real-world financial use case: an autonomous agent that securely accesses market data and executes simulated transactions using financial tools over MCP, demonstrating end-to-end cloud-native deployment.

### 13:45–14:15 · Zero-Trust Traffic Governance For Kubernetes AI Agent Sandboxes

- Room: 1F | Mandarin Hall II
- Speakers: Bingshen Wang, Bo Kang Li
- Track: Networking + Edge + Distributed Systems
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1221856
- Slides: https://sessionize.com/download/ioyjayxe~SGt9nR1NtS54hTts9Bb8dy.pdf~zero-trust_traffic_governance_for_kubernetes_ai_agent_sandboxes.pdf

When AI agents run in Kubernetes, network access can no longer be treated as a static allowlist problem. An agent may execute generated code, choose tools at runtime, call external model or API endpoints, and carry credentials that should not be exposed to the sandbox itself. In our sandbox platform, we treated this as a zero-trust traffic governance problem: every outbound request needs a clear destination, identity, credential boundary, and audit trail.

This talk shares how we built that control path for AI agent sandboxes. We will cover how sandbox traffic is restricted by domain, service, and network target; how platform-wide guardrails and tenant-level allowlists are combined; and how DNS changes, first-packet timing, and short-lived sandboxes affect the design. We will also show how the same path handles identity injection, token replacement, LLM request audit, and forced forwarding through an internal model gateway, without putting a heavy sidecar next to every sandbox.

### 13:45–14:15 · From Chatbots to Agentic AI: Running NVIDIA Dynamo the Kubernetes Way

- Room: 7F | Grand Ballroom II + III
- Speakers: Xianglong Lu, Bo Chu
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Any, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224503

Agentic AI is redefining inference on Kubernetes. Long tool-call workflows create sustained KV-cache pressure; parallel subagents need shared prefixes and consistent routing; and sessions must survive pod/GPU failures without losing context. Prefill–decode disaggregation is becoming a baseline, while multimodal pipelines push architectures toward multi-stage EPD-style decomposition and more stateful coordination.

This talk shares production lessons from running agentic inference on Kubernetes with NVIDIA Dynamo for cache-aware routing and disaggregated serving, and RoleBasedGroup (RBG) for stateful continuity, discovery, and role/unit coordination across pipeline stages. We compare deployment patterns (Dynamo-native vs RBG), what broke under cache contention and partial failures, and how we modeled routing, session continuity, recovery, and observability in Kubernetes-native ways.

### 13:45–15:00 · 📚 Advanced NCCL Tutorial: Beyond Collectives

- Room: 7F | Pearl Hall
- Speakers: Jeff Hammond, Hanyue He, Ke Wen
- Track: 📚 Tutorials
- Labels: AI Infrastructure + Accelerators + Performance Engineering, Advanced, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224567
- Slides: https://sessionize.com/download/okugpa~ErvMRkjRjLFiWksGeRKUuR.pdf~pytorch-china-2026-tutorial.pdf
- GitHub Repository: https://github.com/jeffhammond/pytorch-china-2026-nccl-tutorial

In this tutorial, we will teach you how to use the newest features in NCCL: symmetric memory and GPU-initiated networking (GIN), which are used in state-of-the-art mixture-of-experts (MoE) implementations like DeepEP v2.

This tutorial assumes you know how to initialize NCCL and perform common host-initiated operations. We will explain the concept of symmetric memory and how to use it to scale communication within an NVLink domain. We will also show you how to scale across interconnected GPU nodes using GIN. The examples in the hands-on portion will be simple, but sufficient to demonstrate the concepts and allow attendees to begin working towards complex use cases.

All software materials for this course are open source.

Content for this tutorial will be available in both English and Chinese. The instructors include both native English and Chinese speakers. We will prepare to teach the course in both languages and select based upon the preference of the attendees.

### 13:45–14:15 · Beyond Model Sharding: Atomic Scheduling and Disaggregated LLM Serving with LeaderWorkerSet

- Room: 5F | 5B + C
- Speakers: Kay Yan, Chen Zicong
- Track: Maintainer Track
- Labels: Apps, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222146
- Slides: https://sessionize.com/download/vgayseh~7WqizQY5Mfk3vudj2dh8Bh.pdf~beyond-model-sharding-atomic-scheduling-and-disaggregated-llm-serving-with-leaderworkerset.pdf

LLM serving on Kubernetes is moving beyond model sharding. In production, one inference replica may be a coordinated group of pods across nodes, GPUs, and network topology. If only part of the group is scheduled, accelerators can be reserved without serving traffic. With disaggregated prefill and decode, rollout, scaling, and placement mismatches can also hurt latency and GPU utilization.
LeaderWorkerSet provides a Kubernetes-native API for managing a group of pods as one workload unit for multi-host inference. This talk focuses on two recent areas of LWS evolution: gang scheduling and DisaggregatedSet. We will show how PodGroup integration with Volcano, scheduler-plugins, and YuniKorn schedules leader and worker pods together, and how DisaggregatedSet coordinates multiple LWS resources for prefill and decode.
Attendees will learn when Deployment is no longer enough, when LWS is the right abstraction, and when multiple LWS resources should be managed as one logical service.

### 13:45–15:00 · Peer Group Mentoring

- Room: 5F | 5D + E
- Track: Experiences
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311081

Peer Group Mentoring allows participants to meet with experienced open source veterans across many projects. Mentees are paired with 2 – 8 other people in a pod-like setting to explore technical, community, and career questions together.

If you're interested in attending as a Mentee, seats are first come first served.
If you're interested in being a Mentor, please sign up here: https://www.surveymonkey.com/r/ChinaMentor26

### 14:30–15:00 · Spec-Driven Evolution: How AI Agents are Rewriting the OpenStack/K8s Playbook

- Room: 1F | Mandarin Hall I
- Speakers: Bertrand Souville, Muhammad Hamza
- Track: Networking + Edge + Distributed Systems
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224932

How do you turn complex Telco standards into production-ready code—without months of manual design and validation?

This session introduces a spec-driven development workflow using AI coding agents to accelerate the design, implementation, and testing of features in OpenStack Tacker.
We show how AI agents navigate large codebases to propose architectural solutions and generate test suites directly from specifications.

Two use cases are presented: (1) Cloud-native high availability, with AI-assisted design of Kubernetes Operators enabling redundancy and fast failover toward 99.999% SLA; and (2) Massive-scale vRAN, with a Leader–Conductor style coordination model inspired by OpenStack orchestration systems coordinating thousands of VNFs/CNFs, along with automated edge-case test generation.

Attendees will learn a repeatable workflow for combining AI with upstream collaboration to accelerate the delivery of carrier-grade systems while maintaining OpenInfra and Kubernetes standards.

### 14:30–15:00 · Behind the Scenes of Building a Gateway API Dataplane With Envoy Proxy

- Room: 1F | Mandarin Hall II
- Speakers: Norwin Schnyder, Huabing (Robin) Zhao
- Track: Networking + Edge + Distributed Systems
- Labels: Advanced, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1212758
- Slides: https://sessionize.com/download/plauche~WgfkVQWTwGRhuX3K9a2J4F.pdf~kubecon-china-2026-behind-the-scenes-of-building-a-gateway-api-dataplane-with-envoy-proxy.pdf

Kubernetes Gateway API has become the standard for traffic management in cloud-native environments, with Envoy Proxy emerging as the dominant data plane across implementations. Despite sharing the same proxy, implementations differ in how they translate Gateway API resources into Envoy configuration and distribute it via xDS.

This talk analyzes those differences in depth, covering the trade-offs between State-of-the-World and Delta xDS, how ADS ordering guarantees help achieve hitless configuration rollouts and where they still fall short, and how MuxCache can decouple reconciliation in the controller. We’ll also look at how control planes model backends and routes and what that means for scale, policy attachment, observability, and operational safety.

Drawing on hands-on experience building a Gateway API data plane, you'll leave with lessons learned and actionable insights, whether you're building or operating a Gateway API implementation or evaluating one for production.

### 14:30–15:00 · How Intsig Serves Billions of Document Scans: GPU Virtualization at Scale with HAMi

- Room: 7F | Grand Ballroom II + III
- Speakers: Xiao Zhang, Walter Duan
- Track: AI + ML + Agentic AI + Data Systems
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1221808

Intsig built CamScanner (300M downloads) and runs one of China's busiest document processing platforms. Their GPU problem is unusual: they don't juggle dozens of models — they run one thing (OCR) at absurd concurrency across ~1,000 cards. The bottleneck isn't placement, it's queue time. People waiting in line.

They moved from Tencent QGPU to HAMi, the CNCF Sandbox project, and got virtualization, scheduling, and monitoring in one package. Queue times dropped. Utilization went up.

This isn't just an Intsig talk. HAMi runs in production at SF Express (GPU fleet cut from 1,400 to 1,000 cards), China Merchants Bank (10,000+ GPUs, utilization 20%→80%), NIO (10x CI efficiency), and ICBC (utilization 20%→70%). We'll share patterns that work and anti-patterns that waste money — like why slicing GPUs below 1/6 backfires.

A quick Chaterm demo closes the talk — Intsig's open-source AI terminal for GPU cluster ops in plain language.

### 14:30–15:00 · Dragonfly V2.5.0 - Intro, Updates, Data Distribution in AI Infrastructure

- Room: 5F | 5B + C
- Speakers: Wenbo Qi, Chenyu Zhang
- Track: Maintainer Track
- Labels: Chinese, Dragonfly
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1174542

Dragonfly provides efficient, stable, and secure file distribution and image acceleration using P2P technology within cloud-native architectures. This talk will briefly introduce Dragonfly and highlight the features of its latest version. Key updates include enhanced security and new functionalities tailored for more efficient and robust model distribution. We will also demonstrate how Dragonfly preheats and distributes AI models (packaged as OCI Artifacts) to read-only volumes in Kubernetes, enabling faster deployments. Additionally, we will introduce P2P-based state snapshot and restore capabilities in AI agent scenarios.

### 15:00–15:30 · Coffee Break ☕

- Room: 7F | Grand Ballroom I
- Track: Breaks
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255189

### 15:00–19:00 · Project Pavilion | Tuesday PM Project Tables

- Room: 7F | Grand Ballroom I
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311156

Tuesday PM Project Tables | 15:00 - 19:00

T-1 HAMi
T-2 Volcano
T-3 Argo
T-4 Dragonfly
T-5 CNCF Table
T-6 Community-Led Events
T-7 Cilium
T-8 WasmEdge
T-9 OpenKruise
T-10 K3S
T-11 OpenStack (OpenInfra project)
T-11 Kata Containers (OpenInfra project)
T-12 vLLM (PyTorch project)
T-13 LWS

View the full project table directory here: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/features-add-ons/project-engagement/#project-table-directory

### 15:00–15:30 · Women's Community Gathering (Additional Session)

- Room: 7F | Grand Ballroom I
- Track: Experiences
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1328060

Grab coffee, tea, and a snack during the morning break and meet other women, non-binary, and allies to network, connect and socialize. Look for the Reserved Sign on a table in the Solutions Showcase.

Women’s Gatherings enable meaningful networking, peer mentorship, collaboration, and sustained engagement.

### 15:30–16:00 · Minutes, Not Weeks: Long-Term Memory for AI Applications with PowerMem

- Room: 1F | Mandarin Hall I
- Speakers: Yihang Li
- Track: Application Development + Developer Experience
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1219464
- Slides: https://sessionize.com/download/ilkaywef~wJQphDzZjZhp21FbPgkM5q.pptx~powermem_kubecon_2026_6.0_en.pptx

Every AI app eventually needs memory. Most teams hit the same walls: choosing a backend locks them into a later rewrite, wiring retrieval and lifecycle bleeds into every layer, and non-Python teammates have no clean way in.

PowerMem is an open-source (Apache 2.0) memory library that treats these as developer-experience problems. The talk covers the design choices: auto-loading config from .env so the first example is under ten lines of Python; one memory API across five surfaces (Python SDK, pmem CLI with shell, HTTP API with OpenAPI docs, MCP server over SSE/stdio/streamable-http, web dashboard); pluggable storage (OceanBase, PostgreSQL, SQLite) and model providers so prototype and production share one code path; schema migration and backup/restore baked into the CLI.

Attendees leave with patterns for multi-surface developer experience and a practical checklist for evaluating memory libraries.

### 15:30–16:00 · Building AI4S Systems for Scientific Research on Kubernetes

- Room: 1F | Mandarin Hall II
- Speakers: Jiamu Liu, Yanjun Chen
- Track: Emerging Technologies + Research + Advanced Topics
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224322
- Slides: https://sessionize.com/download/iuytanse~B1n5STzLVTwU3cPDH5uarN.pdf~building-ai4s-systems-for-scientific-research-on-kubernetes0902_v2.0.pdf

Researchers are exploring how AI workflows can accelerate science in recent AI waves. Most existing workflows are built for AI/ML model training, and they fall short for domain-specific scientific work that demands traceable, and reproducible workflows.
To bridge this gap, China Mobile Research Institute built an AI4S system that unifies computing infrastructure and a zero-code software platform. The infrastructure uses Kubernetes to federate multiple cross-region clusters into a pool of heterogeneous resources. The platform covers the entire research lifecycle, form model training to scientific agent development, allowing researchers to build customized assistants without coding.
In collaboration with an RNA-focused biology lab, the system was applied to train a RNA structure prediction model. The resulting model achieved very well performance.
This presentation will share the system architecture, design decisions, and key lessons learned from deploying AI in real scientific research.

### 15:30–16:00 · KernelAgent: Hardware-Guided GPU Kernel Optimization via Multi-Agent Orchestration

- Room: 7F | Grand Ballroom II + III
- Speakers: Kaiming Cheng, Jack Khuu, Laura Wang
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Advanced, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1209151
- Slides: https://sessionize.com/download/shabbek~3ejw5qoAvJxxfrN2Tho1NE.pptx~ptc-china-kernelagent-1.pptx

Recently, the PyTorch team released KernelAgent, an open agentic system achieving 100% correctness across all 250 L1/L2/L3 KernelBench tasks. We extend that work by adding a hardware-guided optimization layer to the existing framework. Building on the previous correctness-focused pipeline, KernelAgent integrates GPU hardware-performance signals into a closed-loop multi-agent workflow to guide the optimization for Triton Kernels.

We evaluate the kernels generated by KernelAgent on all 100 L1 KernelBench tasks. Overall, it achieved 2.02x speedup over generated kernels from earlier versions. On average, KernelAgent generated 1.56x speedup when compared to default torch.compile, outperforming 65 of 100 KernelBench L1 tasks and achieving 89% of the hardware roofline efficiency on the H100.

The optimization codebase is located at KernelAgent repo with documentation to get started. We also share a selection of end-to-end KernelAgent optimization artifacts in the open-source repo.

### 15:30–16:00 · 10 Years of Cilium: Connecting, Securing, and Simplifying the Cloud Native Stack

- Room: 5F | 5B + C
- Speakers: Liyi Huang, Tingjin Ye, Yashi Su
- Track: Maintainer Track
- Labels: Chinese, Cilium
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1227863

Join us as we celebrate a decade of Cilium, now the de-facto standard CNI for Kubernetes and a cornerstone of cloud native networking and security. This session provides updates on the latest Cilium release and highlights how its unified eBPF-powered stack is transforming Kubernetes environments and beyond by replacing fragmented toolchains with seamless, secure, scalable, and simplified solutions.

We’ll showcase how Anta Group used the 'Less is More' approach in their data center architecture by consolidating core networking, load balancing, and observability into Cilium and how Alibaba is using Tetragon to secure AI agents beyond just prompt guardrails. These stories will demonstrate how Cilium is streamlining operations and reshaping the cloud native stack, cementing Cilium’s role as the networking and security data plane for modern infrastructure for the next decade to come.

### 16:15–16:45 · Self-Healing Rollouts: Automating Production Fixes with Agentic AI

- Room: 1F | Mandarin Hall I
- Speakers: Kevin Dubois, Daniel Oh
- Track: Application Development + Developer Experience
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1180062
- Slides: https://sessionize.com/download/lrafcew~phbV87CDwzi9wMN69awkT2.pdf~kubecon-china-self-healing-rollouts_-automating-production-fixes-with-agentic-ai.pdf
- GitHub Repository: https://github.com/kdubois/progressive-delivery
- Project Demo: https://github.com/kdubois/argo-rollouts-quarkus-demo
- Website: https://github.com/kdubois/kubernetes-aiops-agent

Your software rollouts to production are probably always flawless, right? For the rest of us, even with robust CI/CD we do run into occasional production rollout issues. In Kubernetes, Argo Rollouts excels at Progressive Delivery and automated rollbacks, but what if we could go a step further?

This session explores how to elevate your release process by integrating Agentic AI and asynchronous coding agents with canary deployments. We'll demonstrate how an intelligent agent can automatically analyze a rollout failure, pinpointing the root cause. Beyond diagnosis, these agents can take proactive steps on your behalf, suggesting and even implementing code fixes as new PRs, which can be redeployed automatically after PR review. This approach moves us closer to truly self-healing deployments.

Join us to learn how to combine the power of Kubernetes and Argo Rollouts with the autonomous capabilities of Agentic AI, achieving a release experience that is not only seamless but also resilient.

### 16:15–16:45 · Cybertwin-based Cloud Native Network (CCNN): Network Architecture Innovation and Practice

- Room: 1F | Mandarin Hall II
- Speakers: Dandan Liang, Weizhou Lan
- Track: Emerging Technologies + Research + Advanced Topics
- Labels: Advanced, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224457

Driven by IoT and AI, future networks demand an architectural revolution. A CCNN is proposed that converges RAN, IP bearer and data center networks, using Kubernetes as the network OS to unify virtualization, scheduling, and management of compute, storage and network resources, thus enabling secure, cross-domain resource integration. The demo comprises: a Cybertwin-based personal agent service; the security agent function of Cybertwin that implements zero-trust; the data agent function that provides secure multi-dimensional personal data services; the communication agent function that delivers a cloud-native wireless access method for users; finally, a distributed medical collaborative inference based on the user’s personal agent, which demonstrates application-centric unified scheduling and intelligent orchestration of transmission, computing, and storage resources, laying an original, cutting-edge foundation for resource networking and intelligent scheduling in cloud-native networks.

### 16:15–16:45 · Zuul and OpenClaw, Safe Agentic coding via Human Collaboration Tools

- Room: 7F | Grand Ballroom II + III
- Speakers: Monty Taylor
- Track: Application Development + Developer Experience
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1221033

OpenClaw isn't normally used side by side with "safe". But it turns out, for the past 15 years OpenDev and Zuul have been building tools and processes to support safe collaboration between thousands of unknown parties. These tools turn out to be AMAZING at unleashing the potential of systems like OpenClaw, all you have to do is treat your AI agents like untrusted people. What could possibly go wrong?

Using work on the WanderTracks project as an example, we'll walk through the architecture and approach, showing examples of separate long-lived coding and review agents interacting via the same Gerrit review and Zuul gating system we use for everyone else. We'll look at the gerrit and zuul plugins for OpenClaw that allow using code review as communication channels, and we'll highlight the use of Matrix as the primary human to agent communication channel, just like how the Zuul and OpenDev projects. And we'll highlight on the way we're running OpenClaw itself (note, not on my laptop)

### 16:15–16:20 · ⚡ Can Stateful AI Agents Scale to Zero?

- Room: 7F | Pearl Hall
- Speakers: Peng Li
- Track: ⚡Lightning Talks
- Labels: Any, Chinese, Platform Engineering + Cloud Native Architecture
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1215557
- Slides: https://sessionize.com/download/inyascer~rnb26urcM9gvT96VtXLRWJ.pdf~kubecon-can-stateful-ai-agents-scale-to-zero.pdf

AI agents are often treated as always-on workloads because they keep session context, invoke tools, and may depend on sandboxed runtimes. Knative, by contrast, is designed for elastic, scale-to-zero services. This talk shows how the two can work together: by externalizing agent state, separating session continuity from execution, and pooling heavy runtimes, stateful agents can still benefit from serverless elasticity. We will also cover where this model works well—and where always-on runtimes remain the better choice.

### 16:15–16:45 · Evolving Argo Workflows for AI: New Features, Tuning, and Multi-Tenant Isolation

- Room: 5F | 5B + C
- Speakers: Shuangkun Tian, Yashi Su
- Track: Maintainer Track
- Labels: Argo, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1219807

Over the past year, Argo Workflows adoption in AI has surged, particularly in China. Key adopters include BYD, which runs autonomous driving pipelines on Argo, and Qwen3-Coder-Next, which models agentic coding tasks as Argo workflows in its technical report.

As engineers responsible for scaling Argo Workflows in these production environments, we share our insights from operating it as the control plane for large-scale AI pipelines.

We will cover:
• Argo Workflows in Emerging Scenarios: Mapping core capabilities and execution models to autonomous driving, LLM fine-tuning, and agentic coding.
• New Features and Tuning: Recent community enhancements for large-scale workloads, and practical controller tuning patterns for concurrency and queue management.
• Data Isolation and Secure Access: Managing multi-tenant pipelines with per-replica, object-storage-backed volumes, and workflow-aware storage policies so each branch, step, or agent accesses only its own data subset.

### 16:22–16:27 · ⚡Edge Resource Control in KubeEdge:Safely Holding and Releasing Pod Updates for Critical Edge Device

- Room: 7F | Pearl Hall
- Speakers: Feng Gao
- Track: ⚡Lightning Talks
- Labels: English, Intermediate, Networking + Edge + Distributed Systems
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1212199
- Slides: https://sessionize.com/download/ilpavdek~RPVGeztevtqCMdAMGHQwkQ.pptx~edgeresourcecontrolinkubeedge.pptx

In safety-critical edge scenarios like drones in flight, autonomous robots, or AGVs/AMRs, an unplanned pod restart can cause motion interruption, safety hazards, or crashes. Traditional Kubernetes assumes updates can happen anytime, but many edge applications need explicit confirmation that the device is safe (landed, parked, or idle) before upgrading.This lightning talk introduces Edge Resource Upgrade Control, a new feature in KubeEdge v1.22.0. By adding the annotation edge.kubeedge.io/hold-upgrade: "true" to Deployments, StatefulSets, DaemonSets, or Pods, the edge node intercepts the new Pod and holds it in a pending state inside edged. A new HeldUpgrade PodCondition is reported back to the cloud.On the edge, local logic decides when it is safe. Operators release the hold with:keadm ctl unhold-upgrade pod (pod-name)
keadm ctl unhold-upgrade node (node-wide)

### 16:29–16:34 · ⚡ Eliminating Adaptation Latency: Day-0 Inference for New Models on Multi-Backends with vLLM

- Room: 7F | Pearl Hall
- Speakers: Xiaoshuang Wang
- Track: ⚡Lightning Talks
- Labels: Any, Chinese, Hardware Enablement + Diversification
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222625

As LLM iterations accelerate, a significant "adaptation latency"—the days or weeks required to port models from GPU to heterogeneous hardware delays product launches. This bottleneck stems from the lack of clear abstraction boundaries between inference frameworks and hardware backends, forcing invasive code modifications for each new model.
This session decodes how vLLM compresses Ascend NPU adaptation from "per-model porting" to Day-0 Inference through modular design. This session will demonstrate: 1) How pluggable architecture and Out-of-Tree adaptation allow independent development without main-branch blocking; 2) How Custom OP mechanisms enable non-invasive, generalized operator extensions; and 3) How main2main synchronization and robust CI systems ensure Ascend backends stay aligned with vLLM's rapid releases. These capabilities reduce adaptation cycles from weeks to hours and provide a reproducible paradigm for rapid integration within the AI hardware ecosystem.

### 16:36–16:41 · ⚡Empower open edge computing on RISC-V

- Room: 7F | Pearl Hall
- Speakers: Tiejun Chen
- Track: ⚡Lightning Talks
- Labels: Chinese, Intermediate, Networking + Edge + Distributed Systems
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1184985

LF Edge - Linux Foundation Edge, is aiming to establish an open, interoperable software framework for edge computing, independent of hardware, silicon, cloud, but limited on x86 and arm. Meanwhile, open hardware like RISC-V are gaining ground in IoT and embedded systems, with furthering the expansion of edge computing. We believe the blend of such a open edge software, open RISC-V hardware with open LF Edge can herald a new era of empowering Open Software on Open Hardware, to unlock the power of edge computing everywhere.

He we'd like to review how we enable-to-build modern edge platform on real RISC-V hardware platform in the real world by walking through current state, benefits, challenges, and our solutions. Especially, we will demonstrate this with great demo by enabling LF Edge project like EdgeX and edge AI to RISC-V Linux and further show this on heterogeneous hardware platforms along with Linux by deploying on multiple arches as needed.

### 16:43–16:48 · ⚡ Fast Restarts, Not Just Fast Starts: Accelerating Pod Recovery

- Room: 7F | Pearl Hall
- Speakers: Baofa Fan
- Track: ⚡Lightning Talks
- Labels: Any, Application Development + Developer Experience, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1193509
- Slides: https://sessionize.com/download/mvaotte~aXUZycoNMCq6HJsAK45S4o.pptx~fast-restarts.pptx

Pod startup optimization is no longer enough; restart latency is now a core bottleneck for reliability and cost, especially in AI/ML and batch workloads. This session connects start-up and restart as one lifecycle problem. We begin with a practical taxonomy (container restart vs Pod recreation vs rollout/eviction), then break down where restart time is lost: re-scheduling, image pull, init re-execution, and control-plane churn.

The main case study is JobSet in-place restart, where prototype results show recovery improving from 2m10s to 10s at 5,000-node scale. We map this to new Kubernetes capabilities such as per-container restart policies/rules and restart-all-containers. Finally, we provide a production checklist across kubelet tuning, workload controller design, probes/signals, and observability so teams can safely reduce restart time without harming SLOs.

### 16:50–16:55 · ⚡ Speedy But Sturdy PyTorch Optimizers

- Room: 7F | Pearl Hall
- Speakers: Jane Xu
- Track: ⚡Lightning Talks
- Labels: AI Infrastructure + Accelerators + Performance Engineering, Any, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224946
- Slides: https://sessionize.com/download/ttavnew~KeEKMYdVrUMoDuJANVUuYH.pdf~ptc-china-2026-optimizers.pdf

PyTorch strives to offer tested and proven optimizers (think SGD, AdamW, Muon) and we want them to be fast, consistent, and simple to use. We want people to trust our optimizers :D This session will cover how we pushed along each of those dimensions with our latest developments on newer optimizers like Muon, adding more mixed precision (BF16 AdamW), and lifting out-of-the-box performance. We’ll share the new features and algorithms you can try out for driving down your loss curves and, if you’re curious, give insight on how we make things speedier but keep them just as sturdy.

### 16:57–17:02 · ⚡Building a Multi-Cluster Progressive Delivery Platform with Karmada & Argo

- Room: 7F | Pearl Hall
- Speakers: Zhuang Zhang
- Track: ⚡Lightning Talks
- Labels: Any, English, Platform Engineering + Cloud Native Architecture
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1225153

As enterprises scale across hybrid clouds, the gap between "cluster scheduling" and "release control" remains a significant hurdle. While many understand why Karmada and Argo are a powerful duo, the technical "how" of integrating them into a seamless, global delivery pipeline is often left to the imagination.
In this demo-heavy session, we move beyond the theory to show you exactly how to build a unified multi-cluster progressive delivery platform. We will live-demonstrate a complete canary release workflow—from initial configuration to global status aggregation—driven by a single declarative specification.
Key takeaways include:
1. Deep Integration Architecture: How to run ArgoCD and Argo Rollouts with Karmada.
2. Resource Interpretation: Configuring Karmada’s Resource Interpreter to aggregate ReplicaSet status for global observability.
3. Implementation Details: A step-by-step breakdown of the installation, CRD configurations, and propagation policies.

### 17:00–17:30 · Agile Readiness: Leveraging CI Relay and Full-Scale Test Reuse

- Room: 1F | Mandarin Hall I
- Speakers: Jiawei Li, Jingwei Huang
- Track: Hardware Enablement + Diversification
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222276

Adapting to PyTorch’s rapid bi-monthly release cycle is a challenge for heterogeneous accelerators. This session shares Ascend’s practices in maintaining high quality while achieving stable releases within one month of upstream updates.

Test Refactoring & Device-Agnostic Reuse:Learn how we use instantiatedevicetype_tests and dynamic skipping to decouple tests from hardware. This allows seamless reuse of 580K+ existing community test cases, ensuring full API semantic parity and feature validation at minimal cost.

Cross-Repo CI Relay (CRCR) & Quality Gates:We introduce the CRCR mechanism, enabling real-time co-compilation and testing between PyTorch PRs and accelerator code. By intercepting regressions during the PR stage, we mitigate version adaptation pressure and ensure rapid post-release stability.

Through automated test reuse and cross-repo CI, Ascend NPU delivers high-quality releases within 30 days, significantly enhancing user experience.

### 17:00–17:30 · Gravity of Ecosystems: Pollinating the Modern AI Orchestrator

- Room: 1F | Mandarin Hall II
- Speakers: Wenjia Zhang, Bo Fu
- Track: Emerging Technologies + Research + Advanced Topics
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1210983

Open source isn’t just code; it’s the human gravity that keeps our industry moving. However, for too long, infrastructure, orchestration, and AI frameworks have behaved like oil and water—occupying the same space but never truly mixing. This friction creates massive inefficiencies, from "double scheduling" gaps to workloads that remain blind to the hardware they run on.

In this session, Antonio Ojea and Wenjia Zhang share how we are breaking these silos through active pollination. This is not a mere declaration of intent; it is the result of over a year of moving core contributors into projects like Ray, Slurm, and PyTorch and starting dedicated initiatives within Kubernetes. We’ll explore how this hands-on work is evolving the K8s core through Dynamic Resource Allocation (DRA) and the Workload API. By collaborating across foundations instead of working in isolation, we achieve the efficiency and community sovereignty needed to keep the future of AI open and in our hands.

### 17:00–17:30 · Redefining LLM Training through Atomic Components and Serverless TaaS

- Room: 7F | Grand Ballroom II + III
- Speakers: Tan Pei Xiang
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224575
- Slides: https://sessionize.com/download/ciwcacre~GLfFPhvYZigaytCzfBDrC8.pdf~twinkle_deck.pdf

RL (e.g., GRPO) is vital for LLM reasoning, yet current tools are either too rigid or engineering-heavy. Twinkle, co-developed by ModelScope and China Merchants Bank, offers a breakthrough via:
- Atomic Design: Orchestrate RL loops in ~150 lines. Swap modules (Reward/Loss) like Lego without touching the core.
- C-S Architecture: Decouples algorithm semantics from GPU orchestration (torchrun/Ray/HTTP).
- Serverless TaaS: Multi-tenant LoRA pools enable 8+ concurrent tasks on one base model, maximizing resource utility.
As a superset of the Tinker API, Twinkle supports FSDP2 and Megatron backends. It solves the flexibility-scale trade-off, enabling true "Training-as-a-Service." We will explore its architecture, rapid GRPO/DPO integration, and financial sector case studies, concluding with the shift toward "micro-service training" for Agent RL.

### 17:00–17:30 · Hibernate and Evolving: Taming OpenClaw with OpenKruise Agents

- Room: 5F | 5B + C
- Speakers: Zhang Zhen
- Track: Maintainer Track
- Labels: English, OpenKruise
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224605

OpenClaw's rapid adoption has highlighted critical "Day-2" operational challenges: high costs of idle agents, sandbox instability, and complex version upgrades. This session presents a production-ready best practice using OpenKruise Agents to manage these agents efficiently on Kubernetes.

We will demonstrate three core strategies:
1. Secure Execution: Enforcing network isolation and secure API key injection
2. Cost Optimization: Implementing automatic hibernation to scale resources of idle agents to zero, slashing infrastructure costs.
3. Zero-Downtime Upgrades: Evolving OpenClaw versions seamlessly without losing agent state or data.
Finally, we will discuss integrating the upcoming Kubernetes Checkpoint/Restore API to enable vendor-neutral sandbox hibernation. Join us to transform your OpenClaw deployment into a scalable, cost-efficient system.

### 17:04–17:09 · ⚡ Before vLLM starts: Preflight Checks for LWS for LLM Inference on K8S

- Room: 7F | Pearl Hall
- Speakers: Peter Pan
- Track: ⚡Lightning Talks
- Labels: English, Intermediate, Operations + Observability + Reliability
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1220735
- Slides: https://sessionize.com/download/ipcabzec~pCjHasFF5EvMw5K29jsEgj.pptx~kubecon-2026-peter-preflight-and-lws.pptx

Training usually run preflight tests(NCCL, topology, Connectivity) to avoid discovering infra problems after job running.
Inference needs the same discipline. In distributed serving, the most expensive failures are often not startup failures, but clusters that come up and then show unstable throughput, p99 latency jitter, or degraded cross-node performance, with high debugging cost later.

We presents a concrete LWS-based startup gate for distributed inference: run preflight in init-containers, fail early when invariants are broken, and keep checks decoupled from vLLM. As my LWS KEP #813 as the case study (https://github.com/kubernetes-sigs/lws/pull/813), we focus on 3 gaps hit in production: init-phase peer discovery, unbounded recreate loops, and unclear terminal failure boundaries.

To make it practical, we will also share useful examples of preflight checklist, and how to map each failure class to retry, reschedule, or stop.

Fail Early, Inference Better :-)

### 17:30–19:00 · Welcome Reception 🎉

- Room: 7F | Grand Ballroom I
- Track: Experiences
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255206

Join us onsite for drinks, appetizers, and conversations with old and new friends in the Solutions Showcase. Explore the exhibit booths to learn more about the latest technologies, meet experts and project maintainers, browse special offers, and much more.

In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or to access sponsored content. You are never required to visit third-party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.

## Wednesday, September 9, 2026

### 08:00–16:00 · Registration + Badge Pick-Up

- Room: 1F Foyer
- Track: Registration + Badge Pick-up
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255198

### 08:00–17:30 · Cloakroom

- Room: 7F Foyer
- Track: Registration + Badge Pick-up
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255202

### 09:00–09:05 · Welcome Back + Opening Remarks

- Room: 7F | Grand Ballroom II + III
- Speakers: Jonathan Bryce, Horace Li
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1304839

### 09:05–09:15 · What Powers Frontier Intelligence

- Room: 7F | Grand Ballroom II + III
- Speakers: Thierry Carrez
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1312069

Models are becoming more capable, inference more efficient, and cloud native technologies are making AI easier to operate at scale. But everything starts at the infrastructure layer. Every model, inference service and AI agent ultimately depends on the compute, accelerators, networking and storage underneath it.

If you want to build and run the AI stack in open source, the infrastructure layer has to be open too. This keynote explores how OpenInfra makes that vision real—turning heterogeneous hardware and scarce compute into infrastructure that can power AI at scale while providing the isolation, flexibility and control organizations need. From accelerated computing and trusted execution to sovereignty and infrastructure control, we'll look at how open infrastructure connects with cloud native and AI technologies to create a complete open source stack for frontier intelligence.

### 09:17–09:22 · Road from Kata Containers to Confidential Containers + GPUs: From First Commit to CNCF Incubation

- Room: 7F | Grand Ballroom II + III
- Speakers: Zvonko Kaiser
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1304828

Every road has a first step. This one began in November 2021 with a single issue filed against Kata Containers, followed shortly by a first pull request. Talk by talk and release by release, that work led to Confidential Containers becoming a CNCF incubating project with GPU workloads running inside the trust boundary.

This keynote retraces the road through the milestones that shaped it: isolation through lightweight VMs, attestation that reaches from silicon to service, Kubernetes as the platform for AI/ML on private data, and confidential computing on GPUs at scale. Each stop taught a lesson.

Together, they explain where the project stands today and why. The talk ends at incubation, the current end of the road, and closes with a live demo of confidential containers with GPUs.

### 09:24–09:29 · HyperParallel: A SuperPoD-Aware Distributed Acceleration Library

- Room: 7F | Grand Ballroom II + III
- Speakers: Teng Su, Shendi Wang
- Track: Keynote Sessions
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1313956

HyperParallel is a PyTorch-based distributed acceleration library optimized for Ascend SuperPoD. This demonstration is conducted on a Huawei Atlas 800T cluster using a Qwen3-30B-A3B training workload. Under identical configurations, HyperParallel FSDP2 + Muon delivers substantially higher training throughput than PyTorch FSDP2 + Muon while maintaining consistent loss convergence. The live demonstration presents the complete distributed model training workflow and dynamically visualizes training throughput and loss convergence curves, providing an intuitive illustration of HyperParallel’s acceleration capabilities for large-model training.

### 09:31–09:36 · Build a Unified Heterogeneous AI Computing Ecosystem for PyTorch

- Room: 7F | Grand Ballroom II + III
- Speakers: Zesheng Zong
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1309268

AI computing hardware faces severe heterogeneity challenges with high adaptation costs and lack of unified standards. Co-chaired by Huawei and Intel, PyTorch TAC Accelerator Integration WG delivers standardized hardware onboarding guidelines, cross-repo CI testing mechanism, generalized device-aware test suites and platform incubation workflows. This keynote shares our completed achievements and future roadmap including refined device-agnostic APIs and expanded multi-backend test matrix. We invite all hardware vendors and developers to co-build an open, efficient heterogeneous computing ecosystem for PyTorch.

### 09:38–09:41 · Serving Qwen at Scale: Multi-Cluster AI Infrastructure on Karmada

- Room: 7F | Grand Ballroom II + III
- Speakers: Jionghang Cai
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1317079

Serving Qwen requires infrastructure at massive scale: training and serving large language models across multiple Kubernetes clusters and vast GPU resources spanning multiple regions. In this keynote, the Qwen infrastructure team shares how they manage this complexity in production — unifying fragmented GPU resources, scheduling AI workloads across clusters, and keeping the platform reliable as the business grows rapidly. The talk walks through the architecture behind Qwen, the hard lessons learned along the way, and how cloud native multi-cluster technologies, including Karmada, help turn a fleet of clusters into one coherent AI platform.

### 09:45–09:53 · Meet the Community Behind the Open Source AI Stack

- Room: 7F | Grand Ballroom II + III
- Speakers: Jane Lyu, Fupan Li, Zesheng Zong, Hiu Yeung, Sunny Chan
- Track: Keynote Sessions
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1309304

No single project, company, or community builds the AI stack alone. Behind the models, frameworks, cloud native technologies, and infrastructure powering AI is a global open source community working across every layer. This session brings together contributors from across the stack to share how collaboration between projects and communities turns individual technologies into an open, interoperable AI ecosystem—and why that collaboration matters even more as AI evolves.

### 09:55–09:58 · Building Frontier AI Infra: SGLang and Miles

- Room: 7F | Grand Ballroom II + III
- Speakers: Ke Bao
- Track: Keynote Sessions
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1325851

A model’s real-world performance gets decided in the serving and post-training layer: how fast it runs, what it costs, and how quickly people can build on it. That’s the layer SGLang and Miles have been pushing hard on, especially as agentic workloads bring new challenges in cache reuse, long-context serving, and large-scale RL. In this talk, we’ll cover key features designed for agentic workloads, including Unified Radix Cache, HiCache, parallelism strategies for long-context serving, and HiSparse, as well as efficient and scalable RL with Miles, built on top of SGLang for high-performance rollouts and fully asynchronous training.

### 10:00–10:05 · A Cloud Native Stack from Bare Metal to Tokens for Large-Scale AI Inference

- Room: 7F | Grand Ballroom II + III
- Speakers: Trong Vinh Nguyen
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1308448
- Slides: https://sessionize.com/download/iwvabbeb~JtLQi6m8kTTM9XEAGaKDgz.pdf~viettel-ai-platform-2026.pdf

We didn't buy all our GPUs/AI chips at once. Years of procurement across different budget cycles left us with a fleet spanning multiple generations-from T4s that still run fine, to L40Ss, A100s, H200s, and the newer B200s-each with different memory capacities, compute profiles, and driver requirements. Managing these as isolated, team-specific bare-metal allocations meant we were sitting on underutilized assets while other teams waited in queue.

This talk is about how we unified that messy reality into a single platform by layering three open-source ecosystems: OIF at the bare-metal layer, CNCF tooling for GPU pooling and multi-tenant slicing, and PyTorch Foundation projects for serving optimization.

The result is a Token-as-a-Service platform that treats every GPU in the fleet (old or new) as a productive contributor. We'll share what it took to get there, what we're still figuring out, and where we see this going as model complexity and multi-agent workloads continue to grow.

### 10:25–10:30 · Closing Remarks

- Room: 7F | Grand Ballroom II + III
- Track: Keynote Sessions
- Labels: English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1304845

### 10:30–11:00 · Coffee Break ☕

- Room: 7F | Grand Ballroom I
- Track: Breaks
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255190

### 10:30–12:45 · Project Pavilion | Wednesday AM Project Tables

- Room: 7F | Grand Ballroom I
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311158

Wednesday AM Project Tables | 10:30 - 12:45

T-1 OpenEverest
T-2 Karmada
T-3 Cilium
T-4 Dragonfly
T-5 CNCF Table
T-6 Community-Led Events
T-7 llm-d
T-8 spiderpool
T-9 Fluid
T-10 hwameistor
T-11 Kata Containers (OpenInfra project)
T-12 DeepSpeed (PyTorch project)
T-13 kubean

View the full project table directory here: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/features-add-ons/project-engagement/#project-table-directory

### 10:30–15:30 · Solutions Showcase

- Room: 7F | Grand Ballroom I
- Track: Solutions Showcase
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255195

### 11:00–11:30 · Serving AI at Massive Scale: The Cloud-Native Inference Plane Behind Huawei Celia

- Room: 1F | Mandarin Hall I
- Speakers: Kevin Wang
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224671

As AI shifts to integrated mobile assistants, serving multimodal AI across a massive device ecosystem demands managing extreme traffic and strict latency. Prefill-Decode (PD) disaggregation is the essential paradigm, but its sharp compute and bandwidth asymmetry fundamentally breaks traditional Kubernetes Workloads, Routers, and HPA.

Experts from Huawei Consumer Business and Kthena will publicly unpack the infrastructure powering Huawei Celia — one of the world's largest mobile AI assistants. Backed by tens of thousands of accelerators, the presentation will detail PD-disaggregated serving at scale.

Leveraging PodGroup and SubGroup gang scheduling, speakers will explore orchestrating the Inference Plane across Workload, Router, and AutoScaler. The session will analyze real traffic deltas (−40% TTFT, +50% throughput, utilization, and cost-per-million-tokens). Attendees will leave knowing how to model PD workloads, route on tokens, and scale Prefill and Decode using the right metrics.

### 11:00–11:30 · Pathless First: Redefining runc Container Security Against ProcFS-Based Attacks

- Room: 1F | Mandarin Hall II
- Speakers: Fubang Li
- Track: Security + Privacy + Trusted Computing
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224844
- Slides: https://sessionize.com/download/powuy~ReFLDezCLCUwpTkPzRqtMo.pdf~kubecon2026-pathless-first.pdf
- GitHub Repository: https://github.com/opencontainers/runc
- Website: https://lifubang.github.io/acm/

ProcFS has been the most persistent attack surface for runc container escapes, accounting for 60% of path-related runc CVEs from 2017 to 2025. Incremental patches and path hardening have repeatedly failed to address the root cause: runc’s reliance on attacker-controllable ambient path resolution.
As a runc maintainer and lead researcher of the Pathless Container Setup (PCS) framework, I will trace the evolution of ProcFS-based attacks from early TOCTOU symlink hijacking to modern bind mount escapes. I will then present the PCS framework’s file descriptor-centric design, which leverages modern Linux pathless mount APIs (fsopen, fsmount) to completely eliminate path resolution risks. This design has been upstreamed into runc and deployed across global public clouds.
Attendees will learn how this approach mitigates all 9 known Path-related CVEs and gain actionable best practices to harden their own container runtime deployments.

### 11:00–11:30 · Escaping the Vendor Trap: A Journey for Migrating Legacy Infrastructure to OpenStack and K8s

- Room: 7F | Grand Ballroom II + III
- Speakers: Trong Vinh Nguyen, Quoc Dat Le
- Track: Cloud Infrastructure + Virtualization + Storage
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1219415
- Slides: https://sessionize.com/download/wkalwem~3NEA7pbsSRiZT12thET6M1.pdf~viettel-escaping-vendor-trap-2026.pdf

In 2018, most of IT and Telco infras of Viettel (largest telco in Vietnam) run on VMware or legacy on-premise solutions. The utilization of physical infra was low, need hundreds of people to operate all the infras, MTTR was nightmare, depends on support tickets from vendors... Recently, VMWare licenses price went high. That's why we have to migrate our infras to open source tech stacks and solutions. OpenStak, Ceph, Kubernetes, Prometheus... are our choices. By the end of 2025, over 80% of our massive infra had been migrated to open source Cloud and Cloud Native solutions.
This talk will be about how we choose the starting point, how we move from here to there, and the way we achieve the current state, how we collaborated with open source community in Vietnam and Asia.

### 11:00–11:30 · In-Place PVC Re-Binding: Zero-Downtime Disk Migration Using Only the Kubernetes API

- Room: 7F | Pearl Hall
- Speakers: Maxim Nazarenko, Nibir Bora
- Track: Cloud Infrastructure + Virtualization + Storage
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224867

Running stateful workloads in Kubernetes is challenging; migrating the storage beneath them is even harder. Migrating hundreds of live storage volumes without downtime might seem nearly impossible. In this session, we present a battle-tested approach for doing exactly that: migrating hundreds of Azure Disk-backed PersistentVolumes from deprecated in-tree drivers to CSI with zero downtime and minimal service disruption. By leveraging a deep understanding of Kubernetes internals, we accomplished this using only standard user-facing Kubernetes APIs without control plane forks, risky hacks, or even custom operators.

We’ll break down the critical mechanics that make it possible, touching upon immutability constraints, claimRef behavior, finalizers, and StorageClass interactions. Through real-world examples from large-scale production data platforms, this talk demonstrates how advanced Kubernetes systems expertise can unlock operational capabilities that many teams assume are out of reach.

### 11:00–11:20 · Accelerating RL with AgentCube: Cloud-Native Multi-Agent Collaboration

- Room: 5F | 5B + C
- Speakers: Zhencheng Lee
- Track: Sponsor Demos
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1304033

Reinforcement learning workloads often require multiple agents to start concurrently, interact with environments, evaluate rewards, and iteratively update policies. However, Pod cold starts, runtime isolation, and lifecycle management can become bottlenecks for large-scale agent collaboration.

This demo shows how AgentCube provides Kubernetes-native runtime infrastructure for RL workflows. An RL controller creates multiple agent sessions through a unified API, while AgentCube rapidly provisions isolated Pod-based sandboxes using warm pools. The agents then collaborate to perform inference, tool invocation, code execution, and result processing before returning data for reward calculation and policy updates.

The demo will highlight fast startup, multi-agent coordination, session routing, elastic scaling, resource reuse, secure isolation, and automatic cleanup. By integrating with Kubernetes and Volcano scheduling, AgentCube enables high-concurrency, short-lived RL agent workloads with a serverless-like experience.

In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.

### 11:25–11:45 · KV Cache: Accelerating AI inference on Intel CPU

- Room: 5F | 5B + C
- Speakers: Bin Yang
- Track: Sponsor Demos
- Labels: Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1303982
- Slides: https://sessionize.com/download/ijzappeg~GghacyHV9Jff6YjjfU19pa.pdf~intel_accelerating_ai_inference_on_intel_cpu.pdf

Serving LLM is bottlenecked by one scarce resource: HBM. The KV Cache grows with every token and every concurrent request, capping context length, throughput, and ultimately the number of users a cluster can serve. Teams respond by buying more GPUs —paying for compute they don't need just to get the HBM they do.

This talk introduces KVShrink, an approach that treats the KV cache as a tiered, compressible asset instead of a fixed HBM cost. KVShrink offloads the KV cache out of HBM into DDR/Storage and compresses it with Intel QAT hardware accelerator - retaining cache that would otherwise be evicted so it can be reused across requests. Higher hit rates mean the GPU skips redundant re-computation, cutting TTFT latency, improving the user experience, and sustaining far higher concurrency.

In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.

### 11:45–12:15 · Training Through Failures: How Meta Keeps 100k-GPU Jobs Alive with Open-Source Fault Tolerance

- Room: 1F | Mandarin Hall I
- Speakers: Tristan Rice, Amir Afzali
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1223450

GPU failures at scale are not exceptions — they are scheduled events. At Meta, training jobs spanning tens of thousands of GPUs encounter hardware faults daily, yet training continues without missing a step. In this talk, we discuss open-source tools that make this possible.

We introduce the fault tolerance APIs in torchcomms, covering NCCL, MCCL and Gloo backends. MCCL is Meta's production fault tolerance collective communication library and we will deep dive into its design and fault tolerance capabilities. We also show how torchcomms uses NCCL's new primitives — ncclShrink, ncclGrow, and ncclCommRevoke — to gracefully shrink or grow communicator membership without deadlocking the cluster.

We'll cover novel checkpointing approaches — including checkpoint-to-peer-memory and checkpoint-to-local-disk — using torchstore as one example of how fast RDMA transfers between GPU peers can replace slow distributed filesystem writes and enable live recovery.

### 11:45–12:15 · All-or-Nothing No More: Taming Kubernetes Impersonation with Granular Controls

- Room: 1F | Mandarin Hall II
- Speakers: Jian Qiu
- Track: Security + Privacy + Trusted Computing
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1218125
- Slides: https://sessionize.com/download/oksagu~sMh8cviXybcy4WZEN3s9K4.pdf~impersonation-talk.pdf

Kubernetes impersonation lets one identity act as another, but today it's all-or-nothing — an impersonator gains the target's full privileges, often exceeding what's actually needed.
KEP-5284 introduces constrained impersonation, allowing clusters to enforce least-privilege identity delegation by restricting both who can be impersonated and what actions the
impersonator may perform. This unlocks safer patterns for real-world use cases: per-node agents like CSI or CNI plugins that need to impersonate a node but are restricted to only their
required actions, and deputy controllers that securely proxy user operations without inheriting full privileges. Now in beta, we'll walk through the newly introduced authorization
verbs, how constrained impersonation is designed to scale, adoption status from early adopters, and practical guidance for migrating from unrestricted impersonation.

### 11:45–12:15 · Secure AI Agent Sandboxing with OpenStack Zun and Kata Container

- Room: 7F | Grand Ballroom II + III
- Speakers: Wenjie Bao
- Track: AI + ML + Agentic AI + Data Systems
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224832
- Slides: https://sessionize.com/download/idbatkeq~wNbLV3hoKvGDdd9EL11Fiq.pdf~secure-ai-agent-sandbox-upload.pdf

AI Agents are moving to production, but executing LLM-generated code securely at scale remains a critical gap for OpenStack-native environments. Rather than bolting on external services or accepting inadequate isolation, this session presents a secure sandbox built natively on OpenStack.
We’ll share how we combined Zun’s orchestration agility with Kata’s hardware-level isolation to run hundreds of ephemeral, untrusted execution environments, and how this security layer integrates with our broader Agent infrastructure for heterogeneous GPU scheduling and dynamic skill discovery.
This talk will walk through our end-to-end flow: how we leverage OpenStack Glance for sandbox image management, Neutron for isolated tenant networking, and Zun for container lifecycle orchestration; how Kata’s lightweight virtualization provides the hardware-level security boundary between untrusted agent code and the host.

### 11:45–12:15 · One Size Fits None: Lightweight Workload-Specialized Confidential Containers

- Room: 7F | Pearl Hall
- Speakers: Kailun Qin
- Track: Cloud Infrastructure + Virtualization + Storage
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1221758
- Slides: https://sessionize.com/download/isgavos~J6rXBd28tT8FKi77Q8rVqS.pdf~lightweight_coco_kailun-qin.pdf

Confidential Containers (CoCo) brought seamless data-in-use protection to K8s via a Pod-as-a-CVM model with in-guest image pulling and attestation. However, this "compatibility-first" approach comes at the cost of increased complexity, bloated guests, and slow boot times - undermining previous efforts to make virtualized containers lightweight and secure enough to be useful for workloads like serverless and agent sandboxes.

Starting with benchmarks, we analyze the hidden costs and security implications of CoCo. We propose that, for single-purpose workloads, specialization is an essential next step toward the infrastructure "pick-three": speed, scale, and strong security. We dive into how tailored stacks - lightweight VMMs and guest OSes (via unikernels or customized Linux) - fit Confidential Computing context to restore VM-grade isolation with minimal overhead and attack surface. Finally, we highlight the challenges of managing and monitoring these specialized runtimes in production.

### 12:15–13:45 · Lunch (Not Provided) 🍜

- Room: Off-Site Location
- Track: Breaks
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255192

### 13:00–15:00 · Maintainer Meet-Up

- Room: 5F | 5D + E
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1294797

The Maintainer Meetup is for CNCF Maintainers to share best practices, dive into contributing processes, and solve common problems across projects.

### 13:15–15:30 · Project Pavilion | Wednesday PM Project Tables

- Room: 7F | Grand Ballroom I
- Track: Project Opportunities
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1311159

Wednesday PM Project Tables | 13:15 - 15:30

T-1 OpenEverest
T-2 Volcano
T-3 Cilium
T-4 Dragonfly
T-5 CNCF Table
T-6 Community-Led Events
T-7 llm-d
T-8 spiderpool
T-9 OpenKruise
T-10 hwameistor
T-11 OpenStack (OpenInfra project)
T-12 Safetensors (PyTorch project)
T-13 kubean

View the full project table directory here: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/features-add-ons/project-engagement/#project-table-directory

### 13:45–14:15 · Turning Fragmented GPU Clusters Into One Elastic Compute Pool

- Room: 1F | Mandarin Hall I
- Speakers: ChongKang Tan
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1224524
- Slides: https://sessionize.com/download/iplagfel~KNgWE9Ba4DTmsabWTzMdB9.pdf~kubecon-shanghai-2026-slide-0909.pdf

As AI infrastructure evolves, fragmented GPU clusters and bursty AI workloads make large-scale multi-cluster operations increasingly difficult.
In this talk, we share how we built an Elastic Compute Pool that federates fragmented GPU capacity across clusters into a unified global scheduling domain.
While Kubernetes Dynamic Resource Allocation (DRA) and Kueue provide a strong foundation for AI scheduling, we found critical gaps in applying them to global-scale GPU pools—particularly around application-level topology awareness, failover-aware placement, and cross-cluster resource modeling.
We will dive into how we extended DRA to address these limitations and how we unified the orchestration of training and inference workloads across hundred-thousand-GPU-scale infrastructure.
Attendees will learn the architectural tradeoffs, scheduler and resource management extensions, and operational lessons required to make Kubernetes federation practical for AI-native compute platforms.

### 13:45–14:15 · Beyond Static Pods: Dynamic GPU Sharing and Low-Latency Model Switching for LLM Inference on K8s

- Room: 1F | Mandarin Hall II
- Speakers: 夕宁 王, Jun Duan
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1218005

Kubernetes revolutionized microservice orchestration, but Model-as-a-Service exposes a fundamental gap, where MaaS must handle thousands of heterogeneous models and extreme request variability on expensive GPU clusters while meeting strict TTFT/TPOT SLOs. At a large scale inference platform, traditional Kubernetes abstractions prove too coarse-grained: Pod-level model switching takes tens of seconds and idle GPUs remain occupied by inactive models.

We present Fast Model Actuation (FMA), an llm-d incubation project decoupling resource declaration from execution via Dual Pods. A server-requesting Pod holds GPU quota while a server-providing Pod runs inference, dynamically accessing the allocated GPU. Combined with vLLM sleep/wake for fast wake-up and a persistent Multi-Instance Launcher that creates inference subprocesses dynamically, switching drops from minutes to seconds, enabling time-division GPU sharing.

We cover the design, architecture, technical implementation and a demo.

### 13:45–14:15 · To Cache or Not to Cache? A Tiered KVCache Storage System for Agent Scenarios

- Room: 7F | Grand Ballroom II + III
- Speakers: Jingbin Zhang, Wang Cong, yun bai
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Advanced, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1223545

Abstract:
KVCache reuse is a critical lever for reducing inference costs, with scenarios like OpenClaw and Code Agents achieving reuse rates exceeding 90%. While existing KVCache storage systems prioritize transmission throughput, they often overlook lifecycle management—specifically eviction and tiering policies—leaving the storage potential underutilized.

In this talk, we will dive deep into Unified Cache Manager's (UCM) lifecycle management framework, highlighting how heuristic-based dynamic access and eviction policies optimize cache utilization and further drive down inference costs.

The presentation contens:
1. Workload Analysis: Data characteristics and retention patterns across OpenClaw, Code Agents, and general chat sessions.
2. Predictive Modeling: Modeling capacity requirements and retention windows across diverse scenarios, models, and hardware architectures.
3. Policy Implementation: heuristic-driven retention polices and performance gains.

### 13:45–14:15 · From Disk Images to Container Images: Deploying bootc to Bare Metal with Kubernetes and Ironic

- Room: 7F | Pearl Hall
- Speakers: Steve Baker
- Track: Cloud Infrastructure + Virtualization + Storage
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1223326
- Slides: https://sessionize.com/download/cigyatne~hp6btQPKPph7mhbUMHzhnP.pdf~pres.pdf

Bare metal provisioning traditionally relies on building and distributing qcow2 disk images. This results in build pipeline complexity, and limits hardware portability. With bootc (image-mode Linux), a container image contains a complete bootable operating system, including kernel, bootloader, etc. It is stored and distributed identically to any application container image, meaning image registries are available for distribution.

This session presents new OpenStack Ironic capabilities that allow Metal3 to provision bare metal servers directly from bootc container images, with no disk image built or distributed at any point.

The talk begins with a grounding in how container images are stored and how this relates to bootc images and OCI artefacts. Three architectural approaches to bare metal image delivery via Metal3 and Ironic are described. This results in a demonstration showing the end-to-end bootc provisioning flow, from Kubernetes custom resource creation to a deployed host.

### 13:45–14:15 · Helion's New CuTe DSL and Pallas Backends and How we Built Them via Agents

- Room: 5F | 5B + C
- Speakers: Dunfan Lu, Jason Ansel
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1293054

Helion is a Python DSL that raises kernel authoring above languages like CUDA and Triton: you describe a kernel at a high level, and a compiler and autotuner generate the device code. This talk introduces two new compiler backends.

The motivation is hardware heterogeneity and performance. Accelerators are diverging, and reaching state-of-the-art performance on each one normally means hand-writing kernels in that vendor's low-level language. The new backends lower a Helion kernel through CuteDSL for recent NVIDIA GPUs and through Pallas for TPUs. The same source targets different hardware without giving up that performance or rewriting per vendor.

The second half covers code generation by LLM agents. Because Helion sits above the languages it compiles to, a kernel takes fewer tokens to express and offers less surface area for mistakes, and the autotuner handles the performance tuning that agents do poorly. We cover where this holds up and where it doesn't.

### 14:30–15:00 · veRL: Extreme Optimization Practices for Ultra-Large-Scale MoE Models

- Room: 1F | Mandarin Hall I
- Speakers: Carson Wu
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222557

Several leading internet companies (including ByteDance, JD.com, etc.) have built core RL training infrastructure based on veRL to support AI businesses. However, in practical deployment, rapid integration of new models, long-sequence training performance, and training stability remain key challenges.

This talk will introduce the latest technology and engineering practices of the veRL framework. You will learn:

1. We have introduced expert parallelism and sequence parallelism partitioning strategies on top of PyTorch's native FSDP solution.
2. Systematic optimizations for long-sequence training performance:
- Asynchronous RL
- Speculative inference acceleration
- Extreme performance training with low precision
3. Training stability:
Training collapse has long plagued the industry. After sustained and systematic efforts, we have finally achieved full alignment between training and inference.

### 14:30–15:00 · End-to-End Observability for LLM Inference: From Token to GPU

- Room: 1F | Mandarin Hall II
- Speakers: Jared Tan, Murphy Chen, Nicole Li
- Track: Operations + Observability + Reliability
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1223860
- Slides: https://sessionize.com/download/ifiymay~Dur7z1Dbqz7DLEgf5TnNjJ.pdf~end-to-end-observability-for-llm-inference-from-token-to-gpu-en.pdf

LLM inference is becoming the most expensive—and hardest to observe—workload on cloud-native clusters. Unlike stateless microservices, responses are streamed over tens of seconds; performance depends on GPU utilization, KV cache, and continuous batching; and token-level metrics like TTFT, TPOT, and ITL break the classic RED/USE model. A typical Prometheus + Trace stack only tells you the request "succeeded"—not why it was slow, or which layer was slow.

This talk shares our production practice of end-to-end observability for LLM inference, spanning the ingress, Inference Gateway, inference engine (vLLM/SGLang), model runtime, and GPU hardware. We show how to unify metrics across layers, correlate traces through streaming requests, and standardize on the OpenTelemetry GenAI semantic conventions. Three real incidents—a P99 spike, a "busy-but-idle" GPU, and a KV cache OOM—demonstrate how this cuts time-to-resolve from hours to minutes, with a reusable open-source toolchain and configs.

### 14:30–15:00 · Redesigning OpenStack i18n for the AI Era: Weblate & Zero-GPU AI

- Room: 7F | Grand Ballroom II + III
- Speakers: Seongsoo Cho, DaGyeong Kim
- Track: Community + Open Source + Getting Started
- Labels: Beginner, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1225189
- Slides: https://sessionize.com/download/icwavver~ayprGQRRaKPe4WD7h6RkmE.pdf~redesigning-openstack-i18n-for-the-ai-era_-weblate-zero-gpu-ai.pdf

"No open source tool lasts forever." With Zanata reaching end-of-life, OpenStack i18n SIG migrated decades of translation artifacts across 50+ languages to Weblate. Rather than a tool replacement, we saw this as a challenge to redefine the future of open-source internationalization in the AI era.

This session shares our journey of transforming a legacy translation system into a resilient, automated, and AI-ready pipeline at scale.

(1) Migration Strategies: Preserving data integrity across 50+ languages during large-scale workflow transitions

(2) Resilient Architecture: Zuul-powered translation sync and robust failure recovery

(3) Zero-GPU AI Integration: Running open-source LLMs on cloud instances (up to 16 cores) for commit-time draft translations—zero GPU dependency.

Through the OpenInfra University Partnership Program, we mentored Korean students to onboard as active contributors, demonstrating how i18n modernization efforts directly drive sustainable ecosystem growth.

### 14:30–15:00 · From Cloud to AI: How Kata Containers 4.0 Reinvents the Sandbox for the Agent Era

- Room: 7F | Pearl Hall
- Speakers: Fupan Li, Xuewei Niu
- Track: Cloud Infrastructure + Virtualization + Storage
- Labels: English, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1220661

AI agents need more than Linux. Kata Containers 4.0 delivers a universal secure sandbox for the agent era. A new Rust-based default runtime replaces Go for superior safety and speed. Multi-hypervisor support — QEMU, Dragonball, Cloud Hypervisor, Firecracker — adapts to any infrastructure. Template-based instant startup achieves sub-100ms sandbox launch, enabling elastic agent scaling at cloud speed. Confidential Computing via Intel TDX, AMD SEV-SNP, Hygon CSV, and DCU protects sensitive workloads at hardware level. And for the first time, Kata breaks beyond Linux: Android, OSWorld, and Windows sandboxes unify under one runtime, powering diverse AI agent environments securely and instantly. The future of sandboxing is multi-world — and it starts with Kata 4.0.

### 14:30–15:00 · Multi-cluster Orchestration System: Karmada Updates and Use Cases

- Room: 5F | 5B + C
- Speakers: Hongcai Ren, Zongqing Li
- Track: Maintainer Track
- Labels: Chinese, Karmada
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222012

Karmada (Kubernetes Armada) is a Kubernetes management system that enables you to run your cloud-native applications across multiple Kubernetes clusters and clouds.

In this presentation, the maintainer of the Karmada project will share:

- A Brief introduction to Karmada.
- New features over the last year
Application Priority Scheduling
Stateful Application Cluster Failover
Federated ResourceQuota Enforcement
Workload Affinity
AI Jobs Scheduling Enhancements
- Real-world case studies
- Overview of the community
- Roadmap
- QA

### 15:00–15:30 · Coffee Break ☕

- Room: 7F | Grand Ballroom I
- Track: Breaks
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1255191

### 15:30–16:00 · vLLM KV Cache Management: From Cache Reuse to Agent Scenario Optimization

- Room: 1F | Mandarin Hall I
- Speakers: Mengqing Cao, 玺源 王
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222763

As LLM Agents proliferate, the conflict between surging KV Cache and limited VRAM in multi-turn dialogues and tool-chaining has become a key bottleneck. This session shares proven vLLM practices to support longer contexts and higher concurrency under resource constraints.
To address this, vLLM leverages two core mechanisms: Prefix Caching for block-hash-based reuse of common task data (prompts, RAG, history), and KV Offloading to maximize utilization via a three-tier fallback system (LRU, CPU offloading, and ARC policies).
We also explore future roadmap: semantic-aware KV compression based on attention scores, priority-based scheduling for Agent task dependencies, and cross-session prefix sharing. Attendees will understand vLLM’s KV Cache trade-offs, master metrics like Prefix Cache Hit Rate, and preview the evolution of cache optimization for Agent-driven workloads.

### 15:30–16:00 · Non-Invasive AI Agent Observability With OBI

- Room: 1F | Mandarin Hall II
- Speakers: Haibin Zhang, Endre Sara
- Track: Operations + Observability + Reliability
- Labels: Any, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1219307
- Slides: https://sessionize.com/download/islawkes~pgcj26mJd8aau8pRoySAGc.pdf~non-invasive-ai-agent-observability-with-obi-kubecon-extended.pdf

AI applications have evolved from simple LLM calls to sophisticated agent workflows involving tool invocations, MCP interactions, vector embeddings, and RAG. A single agent task may trigger dozens of LLM calls and cross-service tool executions, making debugging extremely challenging.
This talk introduces how OpenTelemetry eBPF Instrumentation (OBI) extends observability capabilities for AI agents. By intercepting traffic at the network layer via eBPF, OBI transparently captures agent behavior — including MCP JSON-RPC tool calls, embedding and reranking operations, and multi-turn LLM-tool interactions. We will demonstrate how eBPF detects MCP Streamable HTTP, extracts tool calls from LLM responses (OpenAI, Anthropic, Gemini, and Qwen), and how the resulting telemetry data aligns with OpenTelemetry GenAI semantic conventions. Attendees will learn how zero-instrumentation observability enables comprehensive and transparent monitoring of complex AI agent systems running on Kubernetes.

### 15:30–16:00 · The State of the PyTorch Ecosystem in 2026: Global Trends and China’s Rising Role

- Room: 7F | Grand Ballroom II + III
- Speakers: Wei Wang
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1225138

In 2026, PyTorch stands at the center of a rapidly expanding global AI ecosystem. Through a data-driven analysis of activity across GitHub, Hugging Face, and related open platforms, this talk reveals a clear upward trajectory in contributions from China—spanning code, models, tools, and community engagement. Yet as domestic efforts in AI frameworks and hardware accelerate, the risk of fragmentation grows. We argue that China’s homegrown AI infrastructure should actively converge with mainstream ecosystems like PyTorch, not to replace them, but to co-develop interoperable standards that benefit the entire global community. The path forward lies not in isolated stacks, but in shared foundations.

### 15:30–16:00 · Write Once, Run Anywhere: PyTorch’s Generalization Journey

- Room: 7F | Pearl Hall
- Speakers: Yu Guangye, Eikan Wang
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222331
- Slides: https://sessionize.com/download/iksavpei~hizufp2XJop4TReHNPCj8d.pdf~pytorch-conference-china-2026-write-once-run-anywhere-pytorchs-generalization-journey.pdf

As AI hardware continues to diversify, PyTorch faces a growing challenge: many of its APIs, runtime interface, and test infrastructures remain fragmented and backend-specific, making it difficult for users to write portable code and for developers to support new hardware consistently.
This talk presents our ongoing generalization effort to make PyTorch truly “write once, run anywhere.” We cover three key areas: unifying module-level APIs such as Autocast, Inductor, and Graph to lower the cost of supporting new hardware; introducing torch.accelerator, a unified runtime API for device, stream, and memory management; and restructuring test infrastructure to validate correctness across multiple backends.
Together, these efforts reduce per-backend maintenance burden, improve test coverage uniformity, and establish a scalable foundation for integrating new hardware into PyTorch with significantly less effort and greater confidence.

### 15:30–16:00 · Solving Industrial Challenges with KubeEdge: A Post-Graduation Report

- Room: 5F | 5B + C
- Speakers: Yue Bao, Huan Wei, Yin Ding
- Track: Maintainer Track
- Labels: Chinese, KubeEdge
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1225125

Following its graduation within the CNCF, KubeEdge has solidified its position as the premier platform for extending Kubernetes to the edge. In this session, project maintainers will explore KubeEdge's evolution, offering a deep dive into the core architecture that enables efficient management of edge workloads.
Attendees will gain insights from real-world deployments across diverse sectors, including Smart Cities, Industrial IoT (IIoT), Edge AI, Robotics, and Retail. Beyond success stories, the talk will cover critical technical updates, including the newly introduced Certified KubeEdge conformance test, recent technological advancements, and the latest updates on community governance.

### 16:15–16:45 · vLLM-Helion: SOTA LLM Performance by advanced autotuning and fine-grained dispatching

- Room: 1F | Mandarin Hall I
- Speakers: Sean Chen
- Track: AI Infrastructure + Accelerators + Performance Engineering
- Labels: Any, English
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1214156
- Slides: https://sessionize.com/download/impaldey~Lu9nPHGesXnMGADsrTkByA.pdf~vllm-helion-pytorch-conference-china-2026.pdf
- GitHub Repository: https://github.com/redhat-et/vllm-helion

High-performance LLM inference demands kernels that are not only algorithmically optimal but also precisely tuned to the hardware and workload at hand. We present vLLM's integration with Helion, PyTorch's Python-embedded kernel authoring DSL.

vLLM-Helion is built around two key ideas: (1) an advanced offline autotuning infrastructure that searches the configuration space across a comprehensive set of input shapes representative of real production workloads, and ships the results as pre-tuned platform-specific configs, eliminating runtime tuning overhead entirely; and (2) a fine-grained runtime dispatching mechanism that selects the optimal pre-tuned configuration based on actual input dimensions at each kernel invocation, rather than relying on a single "one-size-fits-all" config per platform.

Together, these techniques enable vLLM to achieve state-of-the-art kernel performance on many supported hardwares.

### 16:15–16:45 · Why Your TTFT Lies: Diagnosing PD-Disaggregated LLM Inference with Minimal Cross-Layer Metrics

- Room: 1F | Mandarin Hall II
- Speakers: Nicole Li, Kebe Liu
- Track: Operations + Observability + Reliability
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1220647
- Slides: https://sessionize.com/download/idyatweu~QMcLoJtAyViBA2cZx3Mwny.pptx~kubecon-china-2026-why-your-ttft-lies.pptx

Why is inference slow? You have metrics from frameworks/engines/GPUs, but they don't tell you what's wrong. Prefill queueing? Decode memory‑bound? KV transfer latency? GPU/network bottleneck?
With PD‑disaggregated architecture, the problem worsens: different stages behave completely differently, yet observability can't distinguish them.
Root cause: Production LLM inference is a multi‑stage system (Gateway, engines, GPUs, network) — but metrics are siloed. PD‑disaggregation amplifies this: each stage has its own performance behavior, yet monitoring can't tell them apart.

We present a unified observability method for PD‑disaggregated systems. Using a few high-impact metrics, we build a direct mapping: metrics → root cause → tuning action:
- PD‑aware view linking gateway, engine internals (KV Cache, batching, scheduler), infra, GPU & network metrics.
- upgrade from *dashboard watching* to interpretable diagnosis → actionable tuning (P:D ratio, batch size, KV policy, RDMA/topology).

### 16:15–16:45 · What We Learned Securing AI Agents at Scale on Multi-Tenant Kubernetes Clusters

- Room: 7F | Grand Ballroom II + III
- Speakers: Pengfei Ni
- Track: AI + ML + Agentic AI + Data Systems
- Labels: Any, Chinese
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1218010
- Slides: https://sessionize.com/download/ilzawiv~4MTg7dwnN4m6TLkGTZpAu7.pdf~what-we-learned-securing-ai-agents-at-scale-on-multi-tenant-kubernetes-clusters.pdf

Giving every engineer a personal AI agent sounds great until you put hundreds of them on a shared Kubernetes cluster. Each agent executes arbitrary code, calls external tools via MCP, holds cloud credentials, and makes network decisions on its own. Standard Kubernetes multi-tenancy (namespace boundaries, RBAC, NetworkPolicy) was designed for cooperative workloads, not for autonomous agents that can act in unpredictable ways.

We built a multi-tenant agent platform serving hundreds of engineers for daily development and oncall incident response. This talk covers the defense-in-depth architecture we arrived at: Kata Containers giving each agent a VM with its own kernel, Agent ID assigning each agent instance a unique cloud identity via OIDC federation so it can only access data scoped to its owner, NetworkPolicy blocking IMDS and lateral movement, ValidatingAdmissionPolicy preventing exec breakouts, and governance guardrails that enforce blast-radius limits with full audit logging.

### 16:15–16:45 · Kubernetes DRA Architecture: Scheduling, Status, and Topology at Scale

- Room: 7F | Pearl Hall
- Speakers: Paco Xu, Kang Zhang
- Track: Emerging Technologies + Research + Advanced Topics
- Labels: Chinese, Intermediate
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1223308
- Slides: https://sessionize.com/download/kqaxyec~47fBae5CA846z2Fj7Xfbwx.pdf~kubernetes-dra-architecture-v1.1.pdf
- GitHub Repository: https://github.com/NVIDIA/k8s-dra-driver-gpu

AI/HPC workloads need more than “two GPUs”: they need GPUs, NICs, NUMA locality, and PCIe/fabric relationships. This session presents Kubernetes Dynamic Resource Allocation (DRA) in v1.36 as one architecture across entry points, resource modeling, scheduler allocation, binding, kubelet preparation, status, health, and access boundaries. One diagram shows how the main DRA KEPs(14 DRA-related KEPs were updated in v1.36 release cycle) fit together.

The focus is topology-aware scheduling for supernode systems, where NUMA, PCIe, NVLink, or fabric placement determines performance. We will also present an IMEX daemon/channel assignment case study: the design evolved from distributed status updates with conflicts, to centralized updates, and finally to distributed updates on a topology-aligned state object, reducing assignment latency from minutes to seconds at thousand-device scale.

### 16:15–16:45 · Volcano: A Unified Scheduling Platform for Cloud Native AI

- Room: 5F | 5B + C
- Speakers: Chen Zicong, DongYang Wang
- Track: Maintainer Track
- Labels: Chinese, Volcano
- Link: https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1222060

AI infrastructure is moving beyond a single workload model. Training, inference, and agent workloads often run in separate systems today, making resource sharing, queue management, topology placement, and accelerator abstraction harder to operate.

Volcano is evolving as a unified scheduling platform for the AI lifecycle. Volcano-Global extends training jobs across clusters. Kthena is becoming a production-ready LLM serving layer, with rolling updates, router observability, prefix-cache routing, vLLM/SGLang support, and upcoming autoscaling, session affinity, and token-cost-aware balancing. AgentCube brings agent runtimes and code interpreters into Kubernetes with microVM session routing, warm pools, PicoD execution, and SDKs.

At the infrastructure layer, Volcano is improving gang preemption, sharding, namespace queues, throughput, GPU sharing, and HyperNode topology. This session shares the maintainer roadmap and how these pieces converge into one cloud native AI scheduling stack.

