Creating a Synthesia Digital Twin Reveals the Promise and Unease of AI Avatars

相关模型/供应商: OpenAI 供应商
Creating a Synthesia Digital Twin Reveals the Promise and Unease of AI Avatars

An interactive digital version of Alexandru Voica, Synthesia’s head of corporate affairs, introduced Dominic-Madori Davis to a new use of AI in public relations this summer. Rather than simply generating a pitch, Voica’s avatar could answer routine press questions about the company and its technology. For Davis, who had recently joined a panel discussing AI-written pitches, it represented a much bigger step toward automated communications.

In September, Davis visited Synthesia’s new New York office and accepted an invitation to create her own digital twin. Originally based in the U.K., Synthesia competes in the digital-avatar market with companies including D-ID, HeyGen and Colossyan. It reached a $4 billion valuation earlier in 2026 and said in 2025 that its annual recurring revenue had exceeded $100 million.

Synthesia supplies enterprises with AI-avatar training videos. Its recently introduced Roleplay Sessions product lets employees rehearse activities such as sales pitches with an interactive avatar that responds and evaluates their answers.

Building a digital version of a reporter

Davis had previously felt largely indifferent toward avatars, though she expected them to become a regular feature of online life and had seen people use digital likenesses to produce Instagram content. Creating her own made the technology more personally compelling.

The project was described as Synthesia’s first avatar of this kind for a journalist, and its first for anyone other than Voica. The interactive version was restricted to Davis’s reporting on why venture-backed startups commit more fraud than startups without venture backing. It could discuss why she pursued the article, the research paper behind it and the researchers’ findings.

Inside a small studio at Synthesia’s office, the team took numerous photographs and recorded a two-minute voice sample. Davis explicitly consented to the creation of the avatars. The resulting versions included script-reading personal avatars with and without glasses, plus two interactive avatars with the same appearance options.

The models behind the conversation

After selecting the article, a Synthesia team built the interactive avatar using speech recognition, a language model, speech generation and video generation. The setup included Synthesia’s own voice and video models, although customers can select alternatives from providers such as Cartesia, ElevenLabs, Google and OpenAI. Enterprises can choose their cloud hosting environment or pay Synthesia to host the avatars.

The processing sequence starts by converting a user’s speech into text. An agentic language model interprets that text and can take actions based on it. A text-to-voice model generates the spoken response, while Synthesia’s video model animates the avatar.

The company organizes its offerings into three main categories:

  • A video-creation and distribution platform whose conventional avatars deliver typed scripts.
  • An agentic platform called Sessions, supporting interactive experiences such as surveys and roleplay.
  • An API platform that lets developers combine Synthesia’s video and voice models with other services to create interactive avatars or other products.

Testing the resemblance and limits

The team needed a couple of days to produce Davis’s avatars. She first tested the script-reading version with a short passage about autumn arriving in New York, her favorite season. She found the voice reasonably faithful and was pleased that it did not reproduce the hoarseness present during her recording session.

Friends outside the technology industry found the result both intriguing and unsettling. Their response to the interactive version was more critical: they thought its voice and visual resemblance were less convincing than those of the personal avatar, though still close enough to feel eerie.

Davis described the interactive avatar as deterministic and constrained to its assigned article. Questions about her career before TechCrunch or where she lived in New York prompted it to return to the venture-fraud reporting rather than provide personal answers.

Her mother was impressed. Both parents tried asking questions whose answers would be known only within the family, but the avatar consistently redirected them to the article. Its resemblance inspired family humor without overcoming those boundaries.

What avatars could mean for journalism

The experiment prompted Davis to consider whether audiences would accept an avatar delivering the news, whether digital likenesses might supplement or replace journalists, and whether executives would agree to interviews with an AI version of a reporter. One investor immediately rejected the idea of avatar-presented news; others were less certain. Existing resistance to low-quality AI-generated material on social and news-sharing platforms provides an important backdrop to those questions.

For Davis, journalism’s appeal lies in human connections, writing and research. Its central requirement is trust, which she doubts can be delegated to AI.

Outside journalism, she sees potential appeal in a digital counterpart that remains available to answer work questions while its human counterpart is on holiday. How that use develops across corporate America remains uncertain.

Her own enthusiasm became more complicated after the novelty faded. Watching the silent avatar, she found herself waiting for a blink, an unexpected remark or some other sign of awareness. The constrained digital twins she tested would not spontaneously supply that kind of response.

She also expressed concern that a less restricted, freely conversational chatbot behind an avatar could encourage unhealthy psychological attachment or experiences she associated with AI psychosis. This was a personal concern arising from the test, rather than a demonstrated outcome.

Davis predicted that her generation, Gen Z, might struggle to become comfortable with digital twins despite their arrival from what once seemed like science fiction. Still, she found avatars less disconcerting than humanoid robots: a digital encounter can end by logging off. She concluded the experiment by presenting her script-reading avatar delivering a roundup of the site’s leading stories that week.

Dominic-Madori Davis

Dominic-Madori Davis is a New York City-based senior reporter covering venture capital and startups.

分享这篇文章