Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.


(I work at OpenAI.)

It's really how it works.


> (I work at OpenAI.)

Winner of the 'understatement of the week' award (and it's only Monday).

Also top contender in the 'technically correct' category.


> Winner of the 'understatement of the week' award (and it's only Monday).

Yes! As soon as I saw gdb I was like "that can't be Greg", but sure enough, that's him.


and was briefly untrue for like 2 days


[flagged]


Bro what?


How far are we away from something like a helmet with chat GPT and a video camera installed, I imagine this will be awesome for low vision people. Imagine having a guide tell you how to walk to the grocery store, and help you grocery shop without an assistant. Of course you have tons of liability issues here, but this is very impressive


We're planning on getting a phone-carrying lanyard and she will just carry her phone around her neck with Be My Eyes^0 looking out the rear camera, pointed outward. She's DeafBlind, so it'll be bluetoothed to her hearing aids, and she can interact with the world through the conversational AI.

I helped her access the video from the presentation, and it brought her to tears. Now, she can play guitar, and the AI and her can write songs and sing them together.

This is a big day in the lives of a lot of people whom aren't normally part of the conversation. As of today, they are.

0: https://www.bemyeyes.com/


That's definitely cool!

Eventually it would be better for these models to run locally from a security point if view, but this is a great first step.


Absolutely. We're looking forward to Apple's announcements at WWDC this year, which analysts predict are right up that alley.


It sounds like the system that Marshall Brain envisioned in his novella, Manna.


That story has always been completely reasonable and plausible to me. Incredible foresight. I guess I should start a midlevel management voice automation company.


Definitely heading there: https://marshallbrain.com/manna "With half of the jobs eliminated by robots, what happens to all the people who are out of work? The book Manna explores the possibilities and shows two contrasting outcomes, one filled with great hope and the other filled with misery."

And here are some ideas I put together around 2010 on how to deal with the socio-economic fallout from AI and other advanced technology: https://pdfernhout.net/beyond-a-jobless-recovery-knol.html "This article explores the issue of a "Jobless Recovery" mainly from a heterodox economic perspective. It emphasizes the implications of ideas by Marshall Brain and others that improvements in robotics, automation, design, and voluntary social networks are fundamentally changing the structure of the economic landscape. It outlines towards the end four major alternatives to mainstream economic practice (a basic income, a gift economy, stronger local subsistence economies, and resource-based planning). These alternatives could be used in combination to address what, even as far back as 1964, has been described as a breaking "income-through-jobs link". This link between jobs and income is breaking because of the declining value of most paid human labor relative to capital investments in automation and better design. Or, as is now the case, the value of paid human labor like at some newspapers or universities is also declining relative to the output of voluntary social networks such as for digital content production (like represented by this document). It is suggested that we will need to fundamentally reevaluate our economic theories and practices to adjust to these new realities emerging from exponential trends in technology and society."

And a related YouTube video: "The Richest Man in the World: A parable about structural unemployment and a basic income" https://www.youtube.com/watch?v=p14bAe6AzhA "A parable about robotics, abundance, technological change, unemployment, happiness, and a basic income."

My sig is about the deeper issue here though: "The biggest challenge of the 21st century is the irony of technologies of abundance in the hands of those still thinking in terms of scarcity."


Your last quote also reminds me this may be true for everything else, especially our diets.

Technology has leapfrogged nature and our consumption patterns have not caught up to modern abundance. Scott Galloway recently mentioned this in his OMR speech and speculated that GLP1 drugs (which actually help addiction) will assist in bringing our biological impulses more inline with current reality.


Indeed, they are related. A 2006 book on eating healthier called "The Pleasure Trap: Mastering the Hidden Force that Undermines Health & Happiness" by Douglas J. Lisle and Alan Goldhamer helped me see that connection (so, actually going the other way at first). And a later book from 2010 called "Supernormal Stimuli: How Primal Urges Overran Their Evolutionary Purpose" by Deirdre Barrett also expanded that idea beyond food to media and gaming and more. The 2010 essay "The Acceleration of Addictiveness" by Paul Graham also explores those themes. In the 2007 book The Assault on Reason by Al Gore talks about watching television and the orienting response to sudden motion like scene changes.

In short, humans are adapted for a world with a scarcity of salt, refined carbs like sugar, fat, information, sudden motion, and more. But the world most humans live in now has an abundance of those things -- and our previously-adaptive evolved inclinations to stock up on salt/sugar/fat (especially when stressed) or to pay attention to the unusual (a cause of stress) are now working against our physical and mental health in this new environment. Thanks for the reference to a potential anti-addiction substance. Definitely something that deserves more research.

My sig -- informed by the writings of people like Mumford, Einstein, Fuller, Hogan, Le Guinn, Banks, Adams, Pet, and many others -- is making the leap to how that evolutionary-mismatch theme applies to our use of all sorts of technology.

Here is a deeper exploration of that in relation to militarism (and also commercial competition to some extent): https://pdfernhout.net/recognizing-irony-is-a-key-to-transce... "There is a fundamental mismatch between 21st century reality and 20th century security thinking. Those "security" agencies are using those tools of abundance, cooperation, and sharing mainly from a mindset of scarcity, competition, and secrecy. Given the power of 21st century technology as an amplifier (including as weapons of mass destruction), a scarcity-based approach to using such technology ultimately is just making us all insecure. Such powerful technologies of abundance, designed, organized, and used from a mindset of scarcity could well ironically doom us all whether through military robots, nukes, plagues, propaganda, or whatever else... Or alternatively, as Bucky Fuller and others have suggested, we could use such technologies to build a world that is abundant and secure for all. ... The big problem is that all these new war machines and the surrounding infrastructure are created with the tools of abundance. The irony is that these tools of abundance are being wielded by people still obsessed with fighting over scarcity. So, the scarcity-based political mindset driving the military uses the technologies of abundance to create artificial scarcity. That is a tremendously deep irony that remains so far unappreciated by the mainstream."

Conversely, reflecting on this more just now, are we are perhaps evolutionarily adapted to take for granted some things like social connections, being in natural green spaces, getting sunlight, getting enough sleep, or getting physical exercise? These are all things that are in increasingly short supply in the modern world for many people -- but which there may never have been much evolutionary pressure previously to seek out, since they were previously always available.

For example, in the past humans were pretty much always in face-to-face interactions with others of their tribe, so there was no big need to seek that out especially if it meant ignoring the next then-rare new shiny thing. Johann Hari and others write about this loss of regular human face-to-face connection as a major cause of depression.

Stephen Ilardi expands on that in his work, which brings together many of these themes and tries to help people address them to move to better health.

From: https://tlc.ku.edu/ "We were never designed for the sedentary, indoor, sleep-deprived, socially-isolated, fast-food-laden, frenetic pace of modern life. (Stephen Ilardi, PhD)"

GPT-4o, by apparently providing "her" movie-like engaging interactions with an AI avatar that seeks to please the user (while possibly exploiting them) is yet another example of our evolutionary tendencies potentially being used to our detriment. And when our social lives are filled-to-overflowing with "junk" social relationships with AIs, will most people have the inclinations to seek out other real humans if it involves doing perhaps increasingly-uncomfortable-from-disuse actions (like leaving the home or putting down the smartphone)? Not quite the same, but consider: https://en.wikipedia.org/wiki/Hikikomori

Related points by others:

"AI and Trust" https://www.schneier.com/blog/archives/2023/12/ai-and-trust.... "In this talk, I am going to make several arguments. One, that there are two different kinds of trust—interpersonal trust and social trust—and that we regularly confuse them. Two, that the confusion will increase with artificial intelligence. We will make a fundamental category error. We will think of AIs as friends when they’re really just services. Three, that the corporations controlling AI systems will take advantage of our confusion to take advantage of us. They will not be trustworthy. And four, that it is the role of government to create trust in society. And therefore, it is their role to create an environment for trustworthy AI. And that means regulation. Not regulating AI, but regulating the organizations that control and use AI."

"The Expanding Dark Forest and Generative AI - Maggie Appleton" https://youtu.be/VXkDaDDJjoA?t=2098 (in the section on the lack of human relationship potential when interacting with generated content)


I'm more into audiobook, and can't find an audio version for this one. Maybe GPT-4o could read it to me?


Can't wait for the moment when I can puta single line "Help me put this in the cart" on my product and magically sells better.


This Dutch book [1] by Gummbah has the text "Kooptip" imprinted on the cover, which would roughly translate to "Buying recommendation". It worked for me!

[1] https://www.amazon.com/Het-geheim-verdwenen-mysterie-Dutch/d...



Or tell the AI to optimize paper clip production as much as possible.


> Imagine having a guide tell you how to walk to the grocery store

I don't need to imagine that, I've had it for about 8 years. It's OK.

> help you grocery shop without an assistant

Isn't this something you learn as a child? Is that a thing we need automated?


OP specified they were imaging this for low vision people


I'm aware, I'm one of those people.


Does it give you voice instructions based on what it knows or is it actively watching the environment and telling you things like "light is red, car is coming"?


I assume it likes snacks, is quadrupedal, and does not have the proper mouth anatomy or diaphragm for human speech.


Just the ability to distinguish bills would be hugely helpful, although I suppose that's much less of a problem these days with credit cards and digital payment options.



city guide tours, not a bad take tbh :D rather than walkin behind the guy with a megaphone and a flag.


With this capability, how close are y'all to it being able to listen to my pronunciation of a new language (e.g. Italian) and given specific feedback about how to pronounce it like a local?

Seems like these would be similar.


It completely botched teaching someone to say “hello” in Chinese - it used the wrong tones and then incorrectly told them their pronunciation was good.


As for the Mandarin tones, the model might have mixed it up with the tones from a dialect like Cantonese. It’s interesting to discover how much difference a more specific prompt could make.


I don't know if my iOS app is using GPT-4o, but asking it to translate to Cantonese gives you gibberish. It gave me the correct characters, but the Jyutping was completely unrelated. Funny thing is that the model pronounced the incorrect Jyutping plus said the numbers (for the tones) out loud.


Not that different at all.


I think there is too much focus on tones in beginning Chinese. Yes, you should get them right, but no, you'll get better as long as you speak more, even if your tones are wrong at first. So rather than remember how to say fewer words with the right tones, you'll get farther if you can say more words with whatever tones you feel like applying. That "feeling" will just get better over time. Until then, you'll talk as good as a farmer coming in from the country side whose first language isn't mandarin.


I couldn’t disagree more. Everyone can understand some common tourist phrases without tones - and you will probably get a lot of positive feedback from Chinese people. It’s common to view a foreigner making an attempt at Mandarin (even a bad one) as a sign of respect.

But for conversation, you can’t speak Mandarin without using proper tones because you simply won’t be understood.


That really isn't true, or at least it isn't true with some practice. You don't have to consciously think about or learn tones, but you will eventually pick them anyways (tones are learned unconsciously via lots of practice trying to speak and be understood).

You can be perfectly understood if you don't speak broadcast Chinese. There are plenty of heavy accents to deal with anyways. Like Beijing 儿化 or the inability of southerners to pronounce sh very differently from s.


It was good of them to put in example failures.


[flagged]


In my experience, when someone says a project was programmed by "white men from the west coast", it was actually made by Chinese or Indian immigrants.

(Siri's original speech recognition was a combination of Swiss-Germans and people from Boston.)

And it certainly wouldn't be tested by them either way. Companies know how to hire QA contractors.


People always say tech workers are all white guys -- it's such a bizarre delusion, because if you've ever actually seen software engineers at most companies, a majority of them are not white. Not to mention that product/project managers, designers, and QA are all intimately involved in these projects, and in my experience those departments tend to have a much higher ratio of women.

Even beside that though -- it's patently ridiculous to suggest that these devices would perform worse with an Asian man who speaks fluent English and was born in California. Or a white woman from the Bay Area. Or a white man from Massachusetts.

You kind of have a point about tech being the product of the culture in which it was produced, but the needless exaggerated references to gender and race undermine it.


An interesting point, I tend to have better outcomes by using my heavily accented ESL English, than my native pronunciation of my mother tongue I'm guessing it's part of the tech work force being a bit more multicultural than initially thought, or it just being easier to test with

It's a shame, because that means I can use stuff that I can't recommend to people around me

Multilingual UX is an interesting painpoint, I had to change the language of my account to English so I could use some early Bard version, even though It was perfectly able to understand and answer in Spanish


You also get the synchronicity / four minute mile effect egging on other people to excel with specialized models, like Falcon or Qwen did in the wake of the original ChatGPT/Llama excitement.


What? Did it seriously work worse for women? Spurce?

(accents sure)


I don't think that'd work without a dedicated startup behind it.

The first (and imo the main) hurdle is not reproduction, but just learning to hear the correct sounds. If you don't speak Hindi and are a native English speaker, this [1] is a good example. You can only work on nailing those consonants when they become as distinct to your ear as cUp and cAp are in English.

We can get by by falling back to context (it's unlikely someone would ask for a "shit of paper"!), but it's impossible to confidently reproduce the sounds unless they are already completely distinct in our heads/ears.

That's because we think we hear things as they are, but it's an illusion. Cup/cap distinction is as subtle to an Eastern European as Hindi consonants or Mandarin tones are to English speakers, because the set of meaningful sounds distinctions differs between languages. Relearning the phonetic system requires dedicated work (minimal pairs is one option) and learning enough phonetics to have the vocabulary to discuss sounds as they are. It's not enough to just give feedback.

[1]: https://www.youtube.com/watch?v=-I7iUUp-cX8


> but it's impossible to confidently reproduce the sounds unless they are already completely distinct in our heads/ears

interestingly, i think this isn't always true -- i was able to coach my native-spanish-speaking wife to correctly pronounce "v" vs "b" (both are just "b" in spanish, or at least her dialect) before she could hear the difference; later on she was developed the ability to hear it.


I had a similar experience learning Mandarin as a native English speaker in my late 30s. I learned to pronounce the ü sound (which doesn't exist in English) by getting feedback and instruction from a teacher about what mouth shape to use. And then I just memorized which words used it. It was maybe a year later before I started to be able to actually hear it as a distinct sound rather than perceiving it as some other vowel.


After watching the demo, my question isn't about how close it is to helping me learn a language, but about how close it is to being me in another language.

Even styles of thought might be different in other languages, so I don't say that lightly... (stay strong, Sapir-Wharf, stay strong ;)


I was conversing with it in Hinglish (A combination of Hindi and English) which folks in Urban India use and it was pretty on point apart from some use of esoteric hindi words but i think with right prompting we can fix that.


In the "Point and learn Spanish" video, when shown an Apple and a Banana, the AI said they were a Manzana (Apple) and a Pantalón (Pants).


No, I just watched it closely and it definitely said un platano


I re watched it a few times to ensure it said plátano before posting, and it honestly doesn't sound like it to me.


I'm a Spaniard and to my ears it clearly sounds like "Es una manzana y un plátano".

What's strange to me is that, as far as I know, "plátano" is only commonly used in Spain, but the accent of the AI voice didn't sound like it's from Spain. It sounds more like an American who speaks Spanish as a second language, and those folks typically speak some Mexican dialect of Spanish.


> "plátano" is only commonly used in Spain

The wiktionary page for "plátano" has a map illustrating how various Spanish-speaking countries refer to the banana.

https://en.wiktionary.org/wiki/pl%C3%A1tano#/media/File:Porp...

My principal association with plátano is plaintain, personally, but I am not a Spanish speaker.


I was about to comment the same thing about the accent. Even to my gringo ears, it sounds like an American speaking Spanish.

Plátano is commonly used for banana in Mexico, just bought some at a Soriana this weekend.


Interesting, I was reading some comments from Japanese users and they said the Japanese voice sounds like a (very good N1 level) foreigner speaking Japanese.


I thought "plátano" is only used for plantains in Latin America, and Cavendish is typically called "banana" instead. I'm likely wrong, though.


At least IME, and there may be regional or other variations I’m missing, people in México tend to use “plátano” for bananas and “plátano macho” for plantains.


In Spain, it's like that. In Latin America, it was always "plátano," but in the last ten years, I've seen a new "global Latin American Spanish" emerging that uses "banana" for Cavendish, some Mexican slang, etc. I suspect it's because of YouTube and Twitch.


In Spain, plátano is used for Cavendish and plantains are rarely consumed. I am a Spaniard.


I'm from Colombia and mostly say "plátano".


Good to know. I thought Colombians said "banano". That's what a Colombian friend of mine says.


plátano is used in several Spanish-speaking countries, such as Mexico and Chile.


The italian output in the demo was really bad.


I'm a native Italian speaker, it wasn't too bad.


The content was correct but the pronunciation was awful. Now, good enough? For sure, but I would not be able to stand something talking like that all the time


Do you not have to work with non-native speakers of whatever language you use at work?


Most people don't, since you either speak with native speakers or you speak in English mostly, since in international teams you speak in English and not one of the native languages even if nobody speaks English natively. So it is rare to hear broken non-English.

And note that understanding broken language is a skill you have to train. If you aren't used to it then it is impossible to understand what they say. You might not have been in that situation if you are an English speaker since you are so used to broken English, but it happens a lot for others.


Why would you say "really bad"?


It doesn't have hands.


"I Have No Hands But I Must Scream" -Italian Ellison


This was the best joke I’ve heard this year.


So good!


Joke of the day right there :-)


Which video title is this?


Found it in a reel, I’m guessing it’s in the keynote: https://www.instagram.com/reel/C662vlHsGyx/

The Italian sounded good to me.


It sounds like a generic Eastern European who has learned some Italian. The girl in the clip did not sound native Italian either (or she has an accent that I have never heard in my life).


also wondering.


Shared in a reply to my comment.


This is damn near one of the most impressive things, can only imagine especially with live translation and voice synthesis (eleven labs style) you'd be capable of to integrate with something like teams (select each persons language and do realtime translation to each persons native language, with their own voice and intonations would NUTS)


There’s so much pent up collaborative human energy trapped behind language barriers.

Beautiful articulation.

This is an enormous win for humanity.


By humanity you mean Microsoft's shareholders right? Cause for regular people all this crap means is they have to deal with even more spam and scams everywhere they turn. You now have to be paranoid about even answering the phone with your real voice, lest the psychopaths on the other end record it and use it to fool a family member.

Yeah, real win for humanity, and not the psycho AI sycophants


Let AI answer the phone. I love it, can't wait. I hate answering phone but some people just won't email, they always call.


Random OpenAI question: While the GPT models have become ever cheaper, the price for the tts models have stayed in the $15/1Mio char range. I was hoping this would also become cheaper at some point. There're so many apps (e.g. language learning) that quickly become too expensive given these prices. With the GPT-4o voice (which sounds much better than the current TTS or TTS HD endpoint) I thought maybe the prices for TTS would go down. Sadly that hasn't happened. Is that something on the OpenAI agenda?


I've always been wondering what GPT models lack that makes them "query->response" only. I've always tried to get chatbots to lose the initially needed query, with no avail. What would It take to get a GPT model to freely generate tokens in a thought like pattern? I think when I'm alone without query from another human. Why can't they?


> What would It take to get a GPT model to freely generate tokens in a thought like pattern?

That’s fundamentally not how GPT models work, but you can easily build a framework around them that calls them in a loop; you’d need a special system prompt to get anything “thought like” that way, and if you want it to be anything other than stream-of-simulated-consciousness with no relevance to anything, and a non-empty “user” prompt each round, which could be as simple as time, a status update on something in the world, etc.


Monkeys who've trained since birth to use sign language, and can reply incredible questions, have the same issue. The researchers noticed they never once asked a question like "why is the sky blue?" or "why do you dress up". Zero initiating conversation, but they do reply when you ask what they want.

I suppose it would cost even more electricity to have ChatGPT musing alone though, burning through its nvidia cards...


Just provide empty queey and that’s it - it will generate tokens no prob.

You can use any open source model wirthout any promot whatsoever


I think this will be key in a logical proof that statistical generation can never lead to sentience; Penrose will be shown to be correct, at least regarding the computability of consciousness.

You could say, in a sense, that without a human mind to collapse the wave function, the superposition of data in a neural net's weights can never have any meaning.

Even when we build connections between these statistical systems to interact with each other in a way similar to contemplation, they still require a human-created nucleation point on which to root the generation of their ultimate chain of outputs.

I feel like the fact that these models contain so much data has gripped our hardwired obsession for novelty and clouds our perception of their actual capacity to do de novo creation, which I think will be shown to be nil.

An understanding of how LLMs function should probably make this intuitively clear. Even with infinite context and infinite ability to weigh conceptual relations, they would still sit lifeless for all time without some, any, initial input against which they can run their statistics.


It happens sometimes. Just the other day a local TinyLlama instance started asking me questions. The chat memory was full of mostly nonsense and it asked me a completely random and simple question out of the blue. Did chatbots evolve a lot since he was created.

I think you can get models to "think" if you give them a goal in the system prompt, a memory of previous thoughts, and keep invoking them with cron


You might not have a prompt from another human, but you're always receiving new input.


Yes, but that's the fundamental difference. Even if I closed my eyes, plugged my ears and nose and laid in a saltwater floating chamber, my brain will always generate new input / noise.

(GPT) Models toggle between a state of existence when queried and ceasing to exist when not.


You could just let the GPT run in a loop and it too would continue to generate tokens.


And humans malfunction pretty badly without input. Even solitary confinement quickly drives them insane.


> Why can't they?

They are designed for query and reponse. They don't do anything unless you give them input. Also there's not much research on the best architecture for running continuous though loops in the background and how to mix them into the conversational "context". Current LLMs only emulate single thought synthesis based on long-term memory recall (and some goes off to query the Internet).

> I think when I'm alone without query from another human.

You are actually constantly queried, but it's stimulation from your senses. There are also neurons in your brain which fires regularly, like a clock that ticks every second.

Do you want to make a system that thinks without input? Then you need to add hidden stimuli via a non-deterministic random number generator, preferably a quantum based RNG (or it won't be possible to claim the resulting system has free-will). Even a single photon hitting your retina can affect your thoughts and there are no doubt other quantum effects that trips neurons in your brain above the firing threshold.

I think you need at least three of four levels of loops interacting, with varying strength between them. First level would be the interface to the world, the input and output level (video, audio, text). Data from here are high priority and is capable of interrupting lower levels.

The second level would be short term memory and context switching. Conversations needs to be classified, and stored in a database, and you need an API to retrieve old contexts (conversations). You also possibly need context compression (summarization of conversations in case you're about to hit a context window limit).

The third level would be the actual "thinking", a loop that constantly talks to itself to accomplish a goal using the data from all the other levels but mostly driven by the short term memory. Possibly you could go super-human here and spawn multiple worker processes in parallel. You need to classify the memories by asking; do I need more information? where do I find this information? Do I need an algorithm to accomplish a task? What is the completion criteria. Everything here is powered by an algorithm. You would take your data and produce a list of steps that you have to follow to resolves to a conclusion.

Everything you do as a human to resolve a thought can be expressed as a list or tree of steps.

If you've had a conversation with someone and you keep thinking about it afterwards, what has happened is basically that you have spawned a "worker process" that tries to come to a conclusion that satisfies some criteria. Perhaps there was ambiguity in the conversation that you are trying to resolve, or the conversation gave you some chemical stimulation.

The last level would be subconscious noise driven by the RNG, this would filter up with low priority. In the absence of other external stimuli with higher priority, or currently running thought processes, this would drive the spontaneous self-thinking portion (and dreams) when external stimuli is lacking.

Implement this and you will have something more akin to true AGI (whatever that is) on a very basic level.


Train it on stream of consciousness but good luck getting enough training data.


In my ChatGPT app or on the website I can select GPT-4o as a model, but my model doesn't seem to work like the demo. The voice mode is the same as before and the images come from DALLE and ChatGPT doesn't seem to understand or modify them any better than previously.


GPT-4o text version is available not the multi modal one.


I couldn’t quite tell from the announcement, but is there still a separate TTS step, where GPT is generating tones/pitches that are to be used, or is it completely end to end where GPT is generating the output sounds directly?


It's one model with text/audio/image input and output.


Very exciting, would love to read more about how the architecture of the image generation works. Is it still a diffusion model that has been integrated with a transformer somehow, or an entirely new architecture that is not diffusion based?


Licensing the emotion-intoned TTS as a standalone API is something I would look forward to seeing. Not sure how feasible that would be if, as a sibling comment suggested, it bypasses the text-rendering step altogether.


Will the new voice mode allow mixing languages in sentences?

As a language learner, this would be tremendously useful.


Is it possible to use this as a TTS model? I noticed on the announcement post that this is a single model as opposed to a text model being piped to a separate TTS model.


May I just say this launch was a bit of a mess?

The web page implies you can try it immediately. Initially it wasn't available.

A few hours later it was in both the web UI and the mobile app - I got a popu[ telling me that GPT-4o was available. However nothing seems to be any different. I'm not given any option to use video as an input, the app can't seem to pick up any new info from my voice.

I'm left a bit confused as to what I can do that I couldn't do before. I certainly can't seem to recreate much of the stuff from the announcement demos.


The website clearly says that the text version is available now but the multimodal version will be released over the coming weeks.


Who's idea was the singing AIs? What specifically did you want to highlight with that part of the demo?

I imagine that there is a lot of usage at the HQ, human + AI karaoke?


"(I work at OpenAI.)"

Ah yes, also known as being co-founder :)


yes, also known as a programmer loves coding a lot:)


https://community.openai.com/t/when-i-log-in-to-chatgpt-i-am...

Sorry to hijack, but how the hell can I solve this? I have the EXACT SAME error on two iOS devices (native app only — web is fine), but not on Android, Mac, or Windows.


Are you blocking some of your traffic? I had the same issue until I (temporary) disabled NextDNS just for signing in.

Sadly, the error returned is not related to the cause.


No VPN. Mobile Internet and different Wifis. Turned off everything on the devices, from Safari content blockers to IP masking.

Nothing seems too help.


I can't wait to try it out, it sounds too good to be real.

It will be fully available in Eu with the GDPR compliance?


I like the humility in your first statement.


I love how this comment proves the need for audio2audio. I initially read it as sarcastic, but now I can't tell if it's actually sincere.


It’s completely sincere. I’m surprised by the downvotes. Greg Brockman needs no introduction.


> Greg Brockman needs no introduction.

Even if that were true¹, it doesn’t mean everyone would know their HN user name.

¹ Greg may be well known within a select group of people but that’s way smaller than even users of ChatGPT.


I clicked through to see his bio; I didn’t know his username.


And here I thought it was just a GNU debugger fan or something.


debugger, no?


Aye, I have databases on the brain for some reason ... fixed.


I like their username.


You might be talking to GPT-5...


Pretty sure the snark is unnecessary.


I don't think it was snark. The guy is co-founder and cto of OpenAi, and he didn't mention any of that..


I downvoted independently. No problem with groupies. They just contaminate the thread.

Greg Brockman is famous for good reasons but constant "oh wow it's Greg Brockman" are noisy.


Was it snark? To me it sounds like "we all know you Greg"?


This was my intention.


I misunderstood; my apologies.


not snark. if only hn comments could show the right feelings and tonal language


Right to who? To me, the voice sounds like an over enthusiastic podcast interviewer. Whats wrong with wanting computers to sound like what people think computers should sound like?


It sounds VERY California. "Its going great!" "Nice choice" "Whats up with the..." all within 10 seconds.

(not that this is the most important thing about the announcement at all. Just an aside)


It understands tonal language, you can tell it how you want it to talk, I have never seen a model like that before. If you want it to talk like a computer you can tell it to, they did it during the presentation, that is so much better than the old attempts at solving this.


You are a Zoomer sosh meeds influencer, please increase uptalk by 20% and vocal fry by 30%. Please inject slaps, "is dope" and nah and bra into your responses. Throw shade every 11 sentences.


And you’ve just nailed where this is all headed. Each of us will have a personal assistant that we like. I am personally going to have mine talk like Yoda and I will gladly pay Disney for the privilege.


People have been promising this for well over a decade now but the bottleneck is the same as it was before: the voice assistants can't access most functionality users want to use. We don't even have basic text editing yet. The tone of voice just doesn't matter when there's no reason to use it.


I've seen a programmer-turned-streamer literally do this live. Woohoojin on twitch/yt focuses on content for Riot's Valorant esports title, during a couple watch parties he would make "super fans" using GPT with TTS output and the stream of chat messages as input. His system prompts were formed exactly like yours, including instructions to plug his gaming chair sponsor.

It worked surprisingly well. The video where he created the first iteration on stream(don't remember the watch party streams he ran the fans on): https://yewtu.be/watch?v=MBKouvwaru8


I'm not sure whether to laugh or cry...


lowkey genius


Right... enthusiastic and generally confused. It's uncanny valley level expressions. Still better than drab, monotonous speech though.


So far I prefer the neutral tone of Alexa/Google Assistant. I like computers to feel like computers.

It seems like we're in the skeuomorphism phase of AI where tools try to mimic humans like software tried mimic physical objects in the early 2000's.

I can't wait for us to be passed that phase.


Then you can tell it to do that. It will use whatever intonations you prefer.


I want to get to the part where phone recordings stop having slow, full sentences. The correct paradigm for that interface is bullet list, not proper speech.


I can't be the first person that has heard this type of voice before? On a phone tree with a bank when he enter the wrong code?

"It looks like you entered the wrong number! Did you want to try again? Or did you want to talk to an agent?"

That sort of chirpy, overly enthusiastic voice?


> "over enthusiastic podcast interviewer"

Yeh it's cringe. I had to stop listening.

Why did they make the woman sound like she's permanently on the brink of giggling? It's nauseating how overstated her pretentious banter is. Somewhere between condescending nanny and preschool teacher. Like how you might talk to a child who's at risk of crying so you dial up the positive reinforcement.


It's a computer from the valley.


> voice sounds like an over enthusiastic podcast interviewer

I believe it can be toned down using system prompts, which they'll expose in future iterations


As in the Interstellar movie:

    chuckling to 0%

    no acting surprised

    not making bullshit when you don't know


> not making bullshit when you don't know

LLMs today have no concept of epistemology, they don't ever "know" and are always making up bullshit, which usually is more-or-less correct as a side effect of minimizing perplexity.


Oooh, now I want me a TARS...


Genuine People Personalities™, just like in Hitchikers. Perhaps one of the milder forms of 'We Created The Torment Nexus'.


The Total Perspective Vortex in Hitchhiker's notably didn't do anything bad when it was turned on, and so is good evidence that inventing the torment nexus is fine.


What even is this comment

Also,

<spoilers>

It didn't do anything bad to Zaphod Beeblebrox, in a pocket universe created especially for him (therefore ensuring that he was the most important thing in it, and thereby securing his immunity from the mind-scrambling effects of fully comprehending the infinite smallness of one's place in the real universe).


agree I don't get it. I just want the right information and explained well. I don't want to be social with a robot.


exactly. Hope we can customize the voice soon. I want to talk to ultron... or the one from mass effect


>The most impressive part is that the voice uses the right feelings and tonal language during the presentation.

Consequences of audio2audio (rather than audio >text text>audio). Being able to manipulate speech nearly as well as it manipulates text is something else. This will be a revelation for language learning amongst other things. And you can interrupt it freely now!


Anyone who has used elevenlabs for voice generation has found this to be the case. Voice to voice seems like magic.


Elevenlabs isn’t remotely close to how good this voice sounds. I’ve tried to use it extensively before and it just isn’t natural. This voice from openAI and even the one chatGPT has been using is natural.


When have you last used it. I used a few weeks ago to create a fake podcast as a side project recently and it sounded pretty good with their highest end model with cranked up tunings.


About 3 months ago for that exact use case.


My point isn’t necessarily elevenlabs being good or bad, it’s the difference between its text to voice and voice to voice generations. The latter is incredibly expressive and just shows how much is lacking in our ability to encode inflection in text.


However, this looks like it only works with speech - i.e. you can't ask it, "What's the tune I'm humming?" or "Why is my car making this noise?"

I could be wrong but I haven't seen any non-speech demos.


Fwiw, the live demo[0] included different kinds of breathing, and getting feedback on it.

[0]: https://youtu.be/DQacCB9tDaw?t=557


What about the breath analysis?


I did see that, though my interpretation is that breathing is included in its voice tokenizer which helps it understand emotions in speech (the AI can generate breath sounds after all). Other sounds, like bird songs or engine noises, may not work - but I could be wrong.


I suspect that like images and video, their audio system is or will become more general purpose. For example it can generate the sound of coins falling onto a table.


allegedly google assistant can do the "humming" one but i have never gotten it to work. I wish it would because sometimes i have a song stuck in my head that i know is sampled from another song.


I asked it to make a bird noise, instead it told me what a bird sounds like with words. True audio to audio should be able to be any noise, a trombone, traffic, a crashing sea, anything. Maybe there is a better prompt there but it did not seem like it.


The new voice mode has not rolled out yet. It's rolling out to plus users in the next couple weeks.

Also it's possible this is trained on mostly speech.


I was in the audience at the event. The only parts where it seemed to get snagged was hearing the audience reaction as an interruption. Which honestly makes the demo even better. It showed that hey, this is live.

Magic.


I wonder when it will be able to understand that there is more than one human talking to it. It seems like even in today's demo if two people are talking, it can't tell them apart.


I was showing my wife 4o voice chat this afternoon, and we were asking it about local recommendations for breakfast places. All of a sudden…

————

ChatGPT: Enjoy your breakfast and time together.

User: Can you tell that it's not just me talking to you right now?

ChatGPT: I can't always tell directly, but it sounds like you're sharing the conversation with someone else. Is [wife] there with you?

User: My god, the AI has awoken. Yes, this is [wife].

ChatGPT: Hi [wife]! It's great to hear from you. How are you doing?

User: I'm good. Thanks for asking. How are you?

ChatGPT: I'm doing well, thanks! How's everything going with the baby preparations?

—————

We were shocked. It was one of those times where it’s 25% heartwarming and 75% creepy. It was able to do this in part due to the new “memory” feature, that memorized my wife’s name and that we are expecting. it’s a strange novelty now, but this will be totally normalized and ubiquitous quite soon. Interesting times to be living in.


The currently publicly available model uses the old STT -> LLM -> TTS pipeline, so unless you’ve got early access you were using the above.


I'm surprised that ChatGPT is proactively asking questions to you, instead of just giving a response. Is this new? I don't remember this from previous versions.


ChatGPT is now better at being a human than me


How do you know you were talking to 4 O voice, Open AI has not released that yet, just the text version of 4 O


were you using a single account? or is there an setting now to define a family?


It's 90% creepy...


It's already here. Demo: https://youtu.be/kkIAeMqASaY?si=gOspEouo69eQmDiA

I also have an anecdote where it served (successfully) as a mediator for a couple.

Exciting times.


Next announcement: "we've trained our state of the art Situational Awareness Model to solve these problems"


I mention this down thread, but a symptom of a tech product of sufficient advancement is the nature of its introduction matters less and less.

Based on the casual production of these videos, the product must be this good.

https://news.ycombinator.com/item?id=40346002


I noticed this as well. They gave zero fs about the fit and finish of these videos because they know this is magic in a bottle.


That was very impressive, but it doesn't surprise me much given how good the voice mode is in the ChatGPT iPhone app is already.

The new voice mode sounds better, but the current voice mode did also have inflection that made it feel much more natural than most computer voices I've heard before.


Slight off-topic, but I noticed you've updated your llm CLI app to work with the 4o model (plus bunch of other APIs through plugins). Kudos for working extremely fast. I'm really grateful for your tool; I tried many others, but for some reason none clicked as much as your.

Link in case other readers are curious: https://llm.datasette.io


Can you tell the current voice model what feelings and tone it should communicate with? If not it isn't even comparable, being able to control how it reads things is absolutely revolutionary, that is what was missing from using these AI models as voice actors.


No you can't, at least not directly - you can influence the tone it uses a little through the content you ask it to read.

Being able to specifically request different tones is a new and very interesting feature.


+1. Check the demo video in OP titled "Sarcasm". Human asks GPTo to speak "dripping in sarcasm". The tone that comes back is spot on. Comparing that against current voice model is a total sea change.


The voice mode was quite good but the latency and start / stop has been encumbering.


Seems about as good as Azure's Speech Service. I wonder if that's what they are using behind the scenes


"Right" feelings and tonal language? "Right" for what? For whom?

We've already seen how much damage dishonest actors can do by manipulating our text communications with words they don't mean, plans they don't intend to follow through on, and feelings they don't experience. The social media disinfo age has been bad enough.

Are you sure you want a machine which is able to manipulate our emotions on an even more granular and targetted level?

LLMs are still machines, designed and deployed by humans to perform a task. What will we miss if we anthropomorphize the product itself?


This gives me a lot of anxiety but my only recourse is to stop paying attention to AI dev. Its not going to stop, downside be damned. The "We're working super hard to make these things safe" routine from tech companies, who have largely been content to make messes and then not be held accountable in any significant way, rings pretty hollow for me. I don't want to be a doomer but I'm not convinced that the upside is good enough to protect us from the downside.


That's the part that really struck me. I thought it was particularly impressive with the Sal Khan maths tutor demo and the one with BeMyEyes. The comment at the end about the dog was an interesting ad-lib.

The only slightly annoying thing at the moment is they seem hard to interrupt, which is an important mechanism in conversations. But that seems like a solvable problem. They kind of need to be able to interpret body language a bit to spot when the speaker is about to interrupt.


Crazy that interruption also seems to work pretty smoothly


Really? I think interruption and timing in general still seems like a problem that has yet to be solved. It was the most janky aspect of the demos imo.


Yeah, seemed smooth enough but now that I look back it, demo conditions are too perfect which may have made it seem that way.


I’m not sure how revolutionary the style is. It can already mimic many styles of writing. It seems like mimicking a cheerful happy assistant, with associated filler words, etc. is very much in-line with what LLM’s are good at.


Somehow it also sounds almost like Dot Matrix from Spaceballs.


Joan Rivers!


Yeah, the female voice especially is really impressive in the demos. The voice always sounds natural. The male voice I heard wasn't as good. It wasn't terrible, but it had a somewhat robotic feel to it.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: