Companion Apps & Market
miHoYo launched BSide: Olivia Lin on Steam Early Access this month, joining a crowded field. Character AI now reports 233 million registered users, while UnitedHealthcare’s Avery companion is serving 6.5 million members with plans to reach 20.5 million by year-end. Five states have passed or are passing regulatory frameworks around AI companions, reflecting growing mainstream adoption. The market crossed $120 million in annual revenue in April 2026, with platforms growing 38% year-over-year.

Real-Time Video & Avatar Generation
Alibaba’s Wan-Streamer, published in June, is notable as the first native-streaming end-to-end model for audio-visual interaction—it processes language, audio, and video simultaneously in a single model rather than chaining separate modules, enabling true full-duplex conversational video with synchronized facial expressions and lip movements. It remains a research proof-of-concept. On the faster-generation front, ShengShu’s Vidu S1 unveiled real-time interactive video generation, while StreamDiT proposes a streaming architecture achieving 16 FPS on a single GPU. Most production platforms still rely on fast-tier models returning 5–10 second clips in seconds of wall-clock time.

Voice Synthesis
Simba 3.2 ranked first on Artificial Analysis’ TTS leaderboard in July 2026, ahead of ElevenLabs and OpenAI. Fish Audio’s S2 model (open-sourced in March) and Resemble AI’s Chatterbox offer real-time generative audio. By 2026, best-in-class clones reproduce timbre, emotion, and background acoustics well enough that listeners cannot reliably distinguish them over phone or video.

Local LLMs
GLM-5.2 (mid-June from Z.ai) emerged as the strongest all-around open-weight model, with Mistral Large 3 (41B active, 675B total parameters) close behind. Alibaba’s Qwen3 has become the practical default for local deployment, with Apache 2.0 licensing and no commercial restrictions. MiniMax M3 (June 1) combined frontier coding, 1M-token context, and native multimodality in an open-weight model for the first time.

Agentic & Multimodal Reasoning
Agent-X, accepted to ICLR 2026, introduces an 828-task benchmark for vision-centric agentic reasoning spanning images, videos, and mixed-modal instructions across six domains. The benchmark scores every reasoning step and tool use decision; even top models (GPT, Gemini, Qwen) solve fewer than half, exposing meaningful gaps in tool-use and planning capabilities.

Sources:

miHoYo launches AI companion app BSide
Wan-Streamer: Real-Time Multimodal Interaction
ShengShu Vidu S1 Real-Time Interactive Generation
StreamDiT: Streaming Text-to-Video
Best Open Source Voice Cloning Tools
The Best Open-Source LLMs in 2026
Agent-X: ICLR 2026 Benchmark