When Alexandru Voica, head of corporate affairs at video generation startup Synthesia, sent me a link to the latest addition to its PR team this summer, I was surprised. was a interactive virtual avatar of him, trained to answer common press questions about Synthesia, such as what it does and how it works. The day before, I was on a panel where PR people asked me if I cared about proposals that used AI-generated text. But Alexandru’s avatar went beyond that: he seemed to me to be the ultimate boss of the use of AI in public relations.
In September, Synthesia invited me to their new office space in New York. Synthesia, originally based in the United Kingdom, is a popular digital avatar startup along with others like D-ID, HeyGen, and Colossyan. hit a $4 billion valuation earlier this year and said last year that he had crossed $100 million in ARR.
Synthesia allows companies to create interactive training videos with AI avatars and recently launched a product called Role play sessions which allows employees to practice, for example, sales pitches with an interactive AI avatar that responds and rates their answers.
When I went to the opening of the new office and they asked me if I wanted my own AI avatar, I didn’t even hesitate to say yes. Of course I would like to have my own digital twin. My outfit was cute that day and my hair was in place.
Until I met my digital twin, I had been indifferent towards avatars, but I felt that they would inevitably become part of daily life online. I’ve heard of people on Instagram creating them in their image to help them create social content. I find all this very interesting, and perhaps that is why I have no qualms about introducing you to my digital twin now. This is the first time Synethsia has created a digital avatar for a journalist (or anyone, period, outside of Voica). He is trained in my story about why. Venture-backed startups commit more fraud that non-VC-backed startups and will only answer questions about that story.
Next, simply press “start in new window” to get started.
You can ask him questions like:
- Why did I decide to write this story?
- What is the research work about?
- What did the researchers find?
To do this, I entered a mini film studio located inside the Synthesia office where they took numerous photographs of me and captured a two-minute recording of my voice. I had to give my consent for these avatars to be made and, well, digital domination was born. They created a personal avatar for me (one that simply reads any script I give it), with and without glasses, and made two interactive avatars for me (one that can respond to me and listen to me), also with and without glasses.
We choose an item to train the interactive avatar and then one of the equipment I built my interactive avatarwhich works with a combination of speech to text, video, language and text to speech models. My avatar’s tech stack includes Synthesia’s own video and voice models, although the company also allows customers to choose alternatives from other labs such as Cartesia, ElevenLabs, Google or OpenAI. Businesses can also choose to host their avatars on any cloud they want or pay Synthesia to host them.
The speech-to-text model converts what people say into text, the agent language model makes sense of the text and can take actions based on it, the text-to-speech model converts a response into audio, and finally, a video model (created by Synthesia) animates the avatar as it speaks.
In general, Synthesia builds three types of products — a platform for creating and distributing videos with classic avatars, where someone writes a script and the avatar repeats it; an agent platform called Sessions where people can interact with avatars in surveys or Role play; and an API platform where people can take voice and video models from Synthesia and combine them with other technology services to create interactive avatars or other types of products.
It took the Synthesia team a couple of days to create my avatars. I played with the personals first and wrote a fairly generic script to see what my AI voice sounded like. I got him talking about how fall has arrived in New York, my favorite time of year. The voice was quite accurate and I was glad that I didn’t detect the hoarseness I had when I recorded my audio sample.
I showed it to some non-tech savvy friends who found it interesting and creepy.
Then I showed them my interactive, which is deterministic, meaning it will only say what it was trained to respond to. In this case, that was my venture fraud story. I asked him questions like where he was before TechCrunch and what part of New York he lived in, but each time he directed me back to the story.
My friends didn’t think the voice sounded very similar to mine and thought the resemblance wasn’t as good as the personal avatar, but it was still close enough to be kind of creepy. My mom called it “amazing,” which is high praise from her. She and my father kept trying to ask her “only they would know” questions about me, but the model didn’t respond, redirecting them each time to the venture fraud story.
“I don’t remember giving birth to two of you,” she joked after testing the model.
This experience has made me think about what the future of journalism could be. Would people be okay with watching the news and being introduced by an avatar? One investor told me no immediately. certainly there there is a lot of rejection today in AI has infiltrated social networks and other news sharing platforms. But other people I asked weren’t so sure. Could avatars augment (or even replace) journalists? Would CEOs want to talk to an AI avatar of a journalist instead of a human one?
What I like about my career is connecting with people, writing stories, and researching new topics. But the biggest part of journalism is trust. It doesn’t look like you can ever outsource it to an AI.
Outside of journalism, I’m sure the idea of cloning yourself could be appealing. You won’t have to catch up on work after a holiday or vacation, because a version of you can always be there, answering questions.
We’ll have to see how the use of avatars in American companies evolves. But now that I have my own, I have mixed feelings. After getting over the initial novelty, I examined my avatar closely after it stopped talking and waited; I’m not sure what. Maybe I’m waiting for it to blink. Or say something new, or smile, or just let me know that he knows it.
Since my digital twins are deterministic, they will never do that. But I can see how easy it would be for someone to fall into a shade of AI psychosis with one that was non-deterministic, i.e. powered by a chatbot that was free to pontificate.
I told one investor that whatever the future of digital twins holds, I predict that my generation, Generation Z, probably won’t get used to them. They seem like science fiction: everything we have seen in a movie is here now. But I will say that I find digital avatars. less discordant than humanoids. At least with an avatar, if things get weird, I can always log out.
Until then, though, here’s my personal avatar, giving you a rundown of all the top stories on our site this week.
When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.
