Core Event and Technical Positioning

Alipay’s xUI terminal interaction engine, unveiled by platform AI lead Wei Fengdi at AICon 2026 Shenzhen, powers the ‘Abao’ AI entry product. xUI is not a standalone release but a terminal architecture for agentic interaction, covering four core capabilities: full-duplex multimodal communication, generative rendering, service execution, and multi-agent orchestration. Its goal is to transform Alipay’s vast service ecosystem—payments, lifeline services, government offerings, and mini-programs—into conversational Agent actions, addressing the engineering challenge of “user speaks once, Agent finishes the job.”
Technical Evolution and Key Implementations

Communication layer: Early approaches used gRPC for full-duplex and RTC for audio/video, but RTC—optimized for peer-to-peer calls—suffers long connection setup and limited modality extensibility. The team shifted to a three-tier strategy: MoQ as primary (designed for human-Agent interaction), RTC as backup, gRPC as fallback. MoQ natively supports interruption and can multiplex all modalities, dramatically reducing connection latency and network freeze rates.
Rendering layer: Three generations evolved:
- Streaming Markdown rendering: Native rendering (not WebView) ensures ad-hoc display performance and cross-platform protocol support;
- Streaming HTML rendering: In-house Web engine optimizes MY Web and WKWebView, embedding rich charts and interactive components stably;
- A2UI declarative UI: Google-originated framework where Agent generates task steps and engineering composes components. Steps translate via MCP/Skill into UI cards; user confirmation completes the service loop—bridging the gap from"expression" to “completion.”
Service execution layer: For services lacking standard MCP (e.g., government apps, Ant Forest feeds), Alipay combines three approaches:
- GUI Agent: Simulates clicks internally via layout-tree parsing and hit-detection, achieving system-level UI operations without OS permissions;
- TUI Agent: Describes UI structure in text; model infers locations without visual coordinate reasoning—faster but less generic;
- Fixed workflow: For predictable paths (e.g., daily login), skips model inference and runs predefined scripts directly. Critical constraint: execution success rate must exceed 90%; mitigation includes pre-operation page stability verification and cloud-end coordination during interruption.
Cross-Agent Collaboration

Example: User tells phone system assistant “help me collect Ant Energy,” assistant routes to Alipay Agent, which completes the task in background. User need not open Alipay. This avoids cross-app permission breaches and compliance risks, establishing “system assistant handles intent understanding → Alipay handles domain execution”. Partners gain two advantages: background async completion and lower error rates from system-grade capabilities.
Adoption Guidance and Target Audience

- Monitor closely: AI product teams, mobile architects, mini-program developers—A2UI provides a standard protocol path; MCP/Skill integration readiness should be assessed early;
- Consider waiting: Teams expecting plug-and-play GUI automation—GUI Agent generalization for local gov/miniprogram UI is still evolving;
- Architecture reference value: Layered design (network abstraction decouples media and transport) offers reusable patterns for any multimodal application.
Final Thoughts
xUI signifies mobile agents’ shift from dialogue-response to intent-to-action completion. Its engineering value lies not in isolated breakthroughs but in systematic integration—communication, rendering, execution, and orchestration—via protocol decoupling and abstraction, solving the universal challenge of heterogeneous service integration in real-world agent deployment.
