
As a strategic partner of ElevenLabs, Omtera, helps companies move beyond the experimentation stage with AI-powered voice technologies and integrate them into real business processes, customer experiences, and content production operations. One of the most creative features within the ElevenLabs ecosystem, Voice Design, allows you to create a completely new AI voice from scratch simply by describing what you want the voice to sound like, without copying an existing human voice.
Imagine that you need a warm and trustworthy narrator for your brand, an energetic digital assistant for your mobile application, or a calm and easy-to-understand voice for educational content. When you cannot find the voice you are looking for in the existing Voice Library, instead of spending hours searching through different recordings, you can describe the characteristics of the voice you want with a prompt.
Voice Design was developed specifically to solve this problem.
ElevenLabs Voice Design is an ElevenLabs feature that allows users to create a unique synthetic voice by writing a text prompt. Instead of selecting one of the existing voices in the Voice Library, you can create a new voice by describing characteristics such as age, accent, tone, pace, emotion, and speaking style in natural language.
For example, you can define a voice like this:
“A warm and trustworthy female voice in her late 30s. Light British accent, medium speaking pace, natural and professional tone. Clear and calm delivery for corporate training videos.”
Voice Design analyzes this description and generates different voice alternatives that match your specifications.
In the current ElevenLabs Voice Design experience, three different voice previews are generated as a result of each generation. Users can listen to these alternatives, select the most suitable one, and save it to the voice collection in their ElevenLabs account.
This approach is particularly important for teams that want to create a new voice identity tailored to a specific brand personality or use case rather than simply using one of the available ready-made voices.
The Voice Design workflow can essentially be divided into three stages: defining the voice, generating samples, and saving the most suitable voice.
The first step is to determine the purpose for which the voice will be used.
For example:
The intended use of the voice directly affects which details should be prioritized in the prompt.
While a customer support agent may require a trustworthy, clear, and measured voice, a game character may call for a dramatic, unusual, or exaggerated tone.
The primary input of Voice Design is the voice prompt.
According to ElevenLabs documentation, characteristics such as age, gender, accent, tone, pacing, emotion, speaking style, and audio quality can be defined within the prompt. More detailed descriptions can often help the model interpret the intended voice more accurately.
A good Voice Design prompt can include the following elements:
Age: young, middle-aged, elderly, etc.
Voice character: warm, powerful, soft, authoritative, energetic.
Accent: neutral American, British, French, Turkish accent, etc.
Pace: fast, slow, measured, or conversational.
Emotion: calm, excited, trustworthy, friendly.
Use case: audiobook, advertisement, AI agent, or corporate video.
Audio quality: clean studio quality or specific creative effects.
ElevenLabs states that Voice Design v3 also allows elements such as audio quality to be specified through the prompt.
When using Voice Design, it is possible to write a very general description such as “professional male voice.” However, adding context can be a better starting point for more controlled results.
For example, the following prompt could be used for the product promotional videos of a SaaS company:
Voice Design prompt:
Perfect audio quality. A confident male speaker in his late 30s with a neutral international English accent. Warm, professional and trustworthy tone. Medium speaking pace with clear pronunciation and subtle enthusiasm. Designed for enterprise SaaS product videos and executive presentations.
Here, the model learns not only the age and type of voice but also where the voice will be used and what kind of impression it should create for the listener.
A different approach can be used for a customer service AI agent:
A friendly female voice in her early 30s. Calm, patient and reassuring personality. Neutral English accent, natural conversational pacing and clear pronunciation. The voice should sound professional but never robotic, suitable for customer support conversations.
These types of prompts are especially important in Conversational AI projects because one of the first impressions customers form about an agent comes directly from its voice character.
ElevenLabs Voice Design is not simply a tool for “creating a male or female voice.” Many different dimensions that shape a voice character can be defined together.
Characteristics that make a voice sound young, middle-aged, or elderly can be specified within the prompt.
For example:
“Energetic young adult voice”
or
“Calm elderly storyteller with a warm, textured voice.”
Brands can describe specific accent characteristics when producing content for different markets.
For example:
“Neutral American accent”
“Light French accent”
“British English speaker”
can be used.
ElevenLabs states that Voice Design v3 includes improvements for processing more detailed accent combinations and nuanced voice descriptions.
The same text can create a completely different user experience when delivered in different tones.
While a financial application may require a trustworthy and controlled voice, a gaming brand may prefer an energetic and theatrical voice.
Within the prompt, descriptions such as:
can be used.
Voice Design also allows pacing, meaning the speed of speech, to be described within the prompt.
For example:
“slow and reflective”
“fast-paced and energetic”
“measured professional delivery”
can be used.
These two features can sometimes be confused with one another.
Voice Design aims to create a new synthetic voice that does not already exist. The user describes the desired voice in natural language, and the system generates new voice alternatives.
Voice Cloning, on the other hand, is a different process designed to digitally recreate an existing human voice by using its characteristic features.
Therefore, if you want to create a completely new “brand voice” for your company, Voice Design can be a suitable option. If you want to digitally use the voice of a specific and authorized speaker, Voice Cloning serves a different use case.
ElevenLabs also does not position Professional Voice Clone and Voice Design as interchangeable features. In particular, Professional Voice Clone serves a different purpose when there is a need to recreate the voice of a specific licensed speaker.
The process is quite simple within the ElevenLabs web application.
Within ElevenLabs:
Voices → My Voices → Add a new voice → Voice Design
You can follow this path.
Then:
The critical point here is to test the generated voice across different scenarios instead of moving the first result directly into a production environment.
Yes. Voice Design does not have to be used only through the ElevenLabs web interface.
ElevenLabs provides API support for Voice Design. In the official API workflow, voice previews are first generated from a prompt, after which the selected generated voice ID is used to save the voice and use it in other API processes.
This provides a significant advantage, particularly for software teams.
For example, instead of manually creating a voice for every new character, a gaming company could develop a system that automatically generates prompts based on specific character parameters.
Similarly, a content platform could design a workflow such as:
Content brief → Voice persona → Voice Design API → Voice selection → Text to Speech
At this point, a strategic partner with expertise in ElevenLabs such as Omtera can transform Voice Design from a standalone content tool into part of a broader voice AI architecture connected to existing applications, workflows, and enterprise systems.
One of the strongest aspects of Voice Design is that it is not limited to a single industry.
Just as brands standardize colors, fonts, and design systems within their visual identities, they can also standardize their voice identities.
For example, you can create a custom synthetic voice for your company with characteristics such as:
The ability to use custom voices for content production and brand consistency within ElevenCreative is also among the use cases supported by ElevenLabs.
Marketing teams can use Voice Design to create voice characters suitable for different campaigns and formats.
For example, the same company may use:
a more serious voice for a corporate webinar,
a more energetic voice for an Instagram video,
a more explanatory voice for a product demo,
a warmer voice for a customer story.
In Conversational AI systems, the experience is shaped not only by what the agent says, but also by how it says it.
A calm and reassuring voice can be created for a customer support voice agent, while a more energetic yet professional voice may be preferred for a sales agent.
For long educational materials, it is important for the voice to be clear, consistent, and not tiring to listen to.
With Voice Design, a narrator specifically designed for an organization's training content can be created, allowing the same voice character to be used across different modules.
Voice Design v3 is positioned by ElevenLabs for both realistic voice and character voice creation scenarios. Therefore, different character voices can be rapidly prototyped for creative projects ranging from game characters to interactive stories.
Technically, creating a voice may take only a few minutes. However, simply having a “good-sounding voice” is not enough for an enterprise voice that will be used in a production environment.
Before moving to Voice Design, the following questions should be answered:
How does our brand speak?
Who is our target audience?
On which channels will the voice be used?
Should it be formal or friendly?
Energetic or calm?
Which languages will it be used in?
How many minutes will users interact with the voice?
If these questions are not answered before creating the voice, teams may continually experiment with new voices and encounter standardization problems.
Instead of evaluating the voice using only a single demo sentence, it should be tested in real use cases.
For example, if a voice agent will be used, different examples such as:
Greeting message,
long explanation,
question,
apology message,
technical terms,
numbers,
product names
should be tested.
If the voice used by the Marketing team and the voice used by the customer support team have completely different personalities, the brand experience can become fragmented.
Therefore, creating an organization-wide AI voice guideline can be useful.
The guideline can define elements such as:
ElevenLabs' capabilities are not limited to voice generation. The platform enables AI-powered speech generation, Conversational AI, transcription, localization, and API-based voice intelligence scenarios.
Therefore, in an enterprise project, the real value comes not only from generating a new voice but from using that voice within the right business workflow.
As a strategic partner of ElevenLabs, Omtera can support businesses in areas such as:
For example, a company can create a custom voice for customer service using Voice Design. However, in a real production scenario, this voice must be designed together with the Conversational AI agent, knowledge sources, customer data, business logic, and human agent escalation processes when necessary.
This is exactly where the true enterprise value of ElevenLabs technology emerges.
ElevenLabs Voice Design allows companies to create AI voices tailored to their own use cases without being limited to ready-made voice options. From a simple description consisting of a few words to a detailed voice persona containing age, accent, tone, pacing, emotion, and intended use, new voices can be designed with prompts at different levels of detail.
However, especially in enterprise projects, the objective should not be limited to creating an impressive voice. The right voice needs to work together with the right content, workflow, agent, API, and business system. As a strategic partner of ElevenLabs, Omtera can help bring this process to life with the right architecture, from Voice Design to Conversational AI and enterprise voice AI applications.
Ready to create a realistic and scalable AI voice experience? Schedule a quick session with Omtera and bring your ElevenLabs use case to life today.
ElevenLabs Voice Design is a feature that allows users to create a new synthetic AI voice from scratch by writing a text prompt. Characteristics such as age, accent, tone, pace, emotion, and speaking style can be defined.
Within ElevenLabs, go to Voices → My Voices → Add a new voice → Voice Design, describe the voice you want to create, define the preview text, and choose from the generated voice alternatives.
In the current Voice Design experience, three different voice previews are generated during each generation process. Users can select and save the alternative they prefer.
Yes. By defining the voice's age, tone, accent, speaking pace, and personality, you can design a synthetic voice that aligns with the brand's communication style. ElevenLabs also positions custom voice usage as one of the use cases for creating a consistent brand voice.
No. Voice Design creates a new synthetic voice, while Voice Cloning aims to create a digital voice based on the characteristic features of an existing human voice.
Yes. With the ElevenLabs Voice Design API, you can generate voice previews from a prompt, save the selected generated voice ID, and then use it in other ElevenLabs API workflows.
Voice Design can be used for marketing voiceovers, audiobooks, podcasts, game characters, AI voice agents, training content, IVR systems, brand voices, and various digital experiences.
.webp)

