ChatGPT News 05/13/2024 AI Rating: High

OpenAI announces "GPT-4o," unifying text, voice, and vision processing

#ChatGPT#GPT-4o#OpenAI#AI History

On May 13, 2024, OpenAI announced its new flagship model, “GPT-4o.” The “o” stands for “omni,” reflecting its ability to accept and produce a combination of text, voice, images, and video. Voice responses arrive in as little as 232 milliseconds — close to the pace of human conversation — and the fact that it was made available to free users as well drew significant attention.

Details

  • Unified multimodal design: Accepts input combining text, voice, images, and video, and can produce output combining text, voice, and images. Rebuilt as a single model that processes everything end to end
  • Response speed: Responds to voice input in as little as 232 milliseconds, approaching the natural pace of human conversation
  • Performance and cost: Maintains text performance on par with GPT-4 Turbo while achieving twice the speed, half the cost, and 5x higher rate limits via the API
  • Free rollout: Text and image capabilities were made available to free users the same day, while Plus and Team users got a 5x higher message limit
  • Phased rollout: Voice and video capabilities were initially limited to a small group for safety testing, with a phased expansion planned over weeks to months

What happened next

GPT-4o was later released alongside a desktop version of ChatGPT and became the foundation for the natural voice conversation experience delivered by “Advanced Voice Mode.” The policy of extending features to free users carried through into OpenAI’s later product rollouts, and that September led to the announcement of “o1,” a new line of models specialized for reasoning.