AI Companion Products & Market

Tomo AI, an iMessage-native personal coaching and habit-tracking companion, closed a $5M seed round in late June. OpenAI is preparing an AI companion speaker for 2027 — a rechargeable device that will handle cooking assistance, smart home control, and task help. Conversational AI funding remains robust across voice agents, enterprise customer experience, and healthcare workflows.

Real-Time Streaming & Avatar Generation

Wan-Streamer emerged as the first native-streaming end-to-end model combining language, audio, and video in a single system for true full-duplex video calls. Instead of orchestrating separate speech/animation/video modules, it learns cross-modal synchronization as a unified streaming foundation. Wan 2.2, Alibaba’s open-source standard, uses a 27B parameter MoE architecture trained on 1.5B videos.

Live Avatar (ECCV 2026 oral) delivers real-time audio-driven avatar generation at 20 FPS with infinite-length autoregressive support—the model can generate 10,000+ second continuous video. StreamAvatar adapts diffusion models for streaming video generation with low-latency interactive control. Vidu S1 offers real-time voice-controlled character generation at up to 42 FPS on consumer GPUs.

Voice Synthesis

Simba 3.2 ranked top on the independent Artificial Analysis TTS leaderboard this month, outpacing ElevenLabs and OpenAI. The key 2026 advance is prosody modeling: models now predict not just phonemes but delivery timing, stress, and rhythm. Open-source leaders include Fish Speech V1.5, CosyVoice2, and Resemble’s Chatterbox. ElevenLabs v3 (GA February) added multi-speaker dialogue and audio emotion tags.

Local LLMs

Kimi K3 open weights released July 27, 2026, though the full 2.8T MoE model requires cluster compute. GLM-5.2 leads the open-source leaderboard for agentic coding. The market has cleaved into two: frontier MoE models (744B–1.6T parameters) viable only on cloud/enterprise hardware, versus consumer-friendly 4B–31B models (Qwen 3.6, Gemma 4) that fit a single GPU or edge devices.

Agentic & Multimodal Reasoning

Relay-Bench, posted to arXiv in mid-July, chains problems across seven reasoning domains in sequence. GPT-5.5 scores 82.7% on Terminal-Bench 2.0 (CLI agentic performance) but drops to 43.3% on Relay-Bench, exposing a cross-domain reasoning gap. Agent-X (accepted ICLR 2026) benchmarks 828 vision-centric agentic tasks across web, autonomous driving, and math—even top models solve fewer than half. Performance now varies by specialty: Claude leads coding, GPT-5.5 leads agentic breadth, Gemini 3.5 Flash leads speed/multimodal, DeepSeek V4 Pro leads cost-to-performance.

Sources:

Tomo AI seed funding
OpenAI AI companion speaker
Wan-Streamer
Wan 2.2 guide
Live Avatar
StreamAvatar
Vidu S1
Voice cloning 2026 guide
Kimi K3 local setup
Open-source LLM leaderboard
Relay-Bench reasoning benchmark
Agent-X benchmark