The latest AI sound-making technology from OpenAI

OpenAI recently announced limited access to a text-to-voice-producing platform they developed called Voice Engine. This AI sound-making technology is capable of creating synthetic sound based on a person’s 15-second sound clip. The sound produced by this AI can read the text directly in the same language as the speaker or in several other languages. OpenAI said in their blog post, “This small-scale implementation helps clarify our approach, protection, and thinking about how voice engines can be used for the good in various industries.”

company with access

Several companies that have accessed this technology include the education technology company Age of Learning, the Visual Storytelling Heygen platform, the Dimagi frontline health software maker, the creator of the Livox AI communications application, and the Lifespan health system.

Usage example

In the samples posted by OpenAI, we can hear what Age of Learning has done with this technology to produce pre-scripted sound content, as well as read “real-time personal responses” to students written by GPT-4.

audio reference in english

Here are three audio clips produced by AI based on the sample.

OpenAI said they started developing Voice Engine in late 2022 and this technology has been used to turn on pre-set sounds for text-to-voice API and read aloud chat features. In an interview with TechCrunch, Jeff Harris, a member of the OpenAI product team for Voice Engine, said that the model was trained with “a mixture of data licensed and publicly available.” OpenAI told the publication that this model will only be available for about 10 developers.

development of sound generation ai

The AI-to-audio text generation is a thriving area of generative AI. While most focus on instrumental or natural sound, few focus on the generation of sound, in part because of the questions asked by OpenAI. Some of the names in the space include companies such as Podcastle and ElevenLabs, which provide AI sound cloning technology and tools that Vergecast has explored last year. Meanwhile, the US government is trying to control the unethical use of AI sound technology. Last month, the Federal Communications Commission banned spam calls using AI voice after people received spam calls from President Joe Biden’s voice.

According to OpenAI, their partners agree to comply with their use policy which states that they will not use Voice Generation to disguise themselves as people or organizations without their consent. It also requires partners to get “explicit and informed consent” from the original speaker, not building a way for individual users to make their own voice, and to reveal to listeners that the sounds are generated by AI. OpenAI also adds a watermark to the audio clip to track its origin and actively monitors how the audio is used.

Security measures

OpenAI suggests several measures that he said could limit the risks around such tools, including stopping voice-based authentication to access bank accounts, policies to protect the use of AI insiders, further education on AI Deepfakes, and AI content tracking system development.

FAQs Related to the Latest AI Voice Making Technology from OpenAI

  1. What is the voice engine of OpenAI?

    The Voice Engine is the latest AI sound generation platform from OpenAI that enables the creation of synthetic sound based on one’s 15-second sound clip.

  2. Who has accessed Voice Engine technology?

    Several companies that have accessed this technology include Age of Learning, Heygen, Dimagi, Livox, and Lifespan.

  3. How does OpenAI ensure the ethical use of the Voice Engine?

    OpenAI ensures ethical use by seeking explicit approval from the original speaker, prohibiting the use of undercover, and adding a watermark to track the origin of the audio clip.

  4. What sets the voice engine apart from other AI voice generation technologies?

    The Voice Engine distinguishes itself with its ability to produce synthetic sound based on short sound clips and provides strict control over its use.

  5. What are the steps suggested by OpenAI to limit the risks surrounding AI sound generation technology?

    OpenAI suggests several steps, including stopping voice-based authentication to access bank accounts and developing an AI content tracking system to reduce the risk of abuse.

Baca Juga

Back to top button

Adblock Detected

LidahTekno.com is supported by Google Adsense advertising to provide content for you.Please consider disabling AdBlocker or adding us to your whitelist so we can continue providing the best technology information and tips.Thank you for your support!