Trying Talk Machine: first impressions of a voice-first AI app

0.00 avg. rating (0% score) - 0 votes

I recently spent some time testing Talk Machine, a voice-first app available on iOS, Android, and the web. I mainly tested the iOS app, with a brief look at the web version for voice playback. The idea is simple but interesting: instead of typing into a chatbot, you talk to the app, and it replies with both text and speech. It also supports voice notes through email, which makes the product feel less like a closed messaging app and more like a voice layer that can work with existing communication channels.

The app is still early and there are some rough edges, but the main experience is already pleasant. The interface is friendly, the onboarding is quick, and the voice interactions are more natural than I expected. It does not feel like a traditional productivity app with a microphone button added as an afterthought. The whole interface is built around speaking first.

Account setup

The login process is simple. I entered my email address, received a verification code almost instantly, and was able to start using the app without CAPTCHA, payment information, or a long account setup flow. For a voice app, this is a good first impression because the user can get to the actual voice experience quickly.

The app prompted me for access to Gmail and contacts, but the permission was optional. I could deny it and continue using the app, then add people manually later. This is a sensible approach because contact integration is useful, but it should not be required before the user has had a chance to understand the product.

TalkMachine Home Screen

The home screen shows several built-in contacts or bots, including newsbot, helpbot, askbot, timebot, and weatherbot. There is also a contact-us entry and the user profile. The visual design is simple and friendly, with large buttons, rounded cards, and a prominent voice control. It feels more like a small communication device than a conventional chat application.

Voice setup

The voice setup wizard asks the user to read a short sample paragraph. The app then tries to learn what the user’s voice sounds like and plays back a new piece of speech using a similar voice. The generated voice sounded somewhat like me. Not exact, but close enough to make the feature interesting.

Voice Setup Paragraph

The setup paragraph sounds almost like an excerpt from a memoir. It is short enough that most users will probably finish it. Voice products can become tedious if they ask the user to configure too much audio-related information up front. Here, the process is just: read a paragraph, hear the result, continue.

Every evening on my walk home, somewhere between the station and my front door, I send a voice note to my friend.

I tell her about my day, one good thing and one funny thing. The next morning I have her reply and I listen to it over my coffee.

I like memoirs, by the way, so I did not mind the tone. It is also a good paragraph for this purpose because it is conversational and easy to read aloud. A more technical paragraph would probably be less useful for learning the user’s speaking style.

Speech recognition and transcript cleanup

One interesting behavior is that speech recognition appears to happen in more than one pass. Immediately after I finished recording, the app displayed an initial transcript. A short moment later, the text changed into a cleaner version that seemed to use more context and better sentence reconstruction.

In one test, the first version looked like this:
I will go to look at the video and answer the question from there…
Speech Recognition First Pass New

After refinement, it became much closer to what I actually said:
Are you able to look at video and answer questions from there, or at the moment can you only answer from voice?
Speech Recognition Second Pass New

This is a useful behavior because raw speech-to-text output is often awkward. Real speech is messy. People pause, restart sentences, speak in fragments, and sometimes record in noisy environments. A second cleanup pass can make the final transcript much more readable. The only confusing part is that the user may briefly see the first, less accurate transcript before it changes. If this is intentional, a small “refining transcript” or similar status would make the behavior easier to understand.

Voice notes through email

The email-based voice note flow is one of the most interesting parts of Talk Machine. I recorded a message and sent it to another email address.

Voice Chat conversation via email

The recipient received an email containing the transcript, a playable recording, and a button to reply by voice through the web interface.

Voice Chat Msg Received

The email transcript was readable, and the reply-by-voice button was easy to find. This flow could be useful for people who prefer speaking but still need to communicate with others who mostly use email. It also makes the product easier to try because the first recipient does not need to understand the whole app before replying.

Built-in bots

I tried several of the built-in bots, including askbot, helpbot, newsbot, weatherbot, and timebot. The results were mixed, but the general direction is clear. Newsbot handled a request for Vietnamese headlines with English translation without any major issue. I asked for headlines about Vietnam in Vietnamese and also asked it to provide the English version. The response came back in both languages and was easy to read.

Vietnamese Answers

Askbot also gave a reasonable answer when I asked how it was different from ChatGPT. It described itself as being built into Talk Machine, optimized for quick spoken answers, and able to handle reminders and recurring schedules inside the chat.

better than chatgpt

The separate bots make the home screen easy to understand visually, but they also create some boundaries. One bot may refuse to answer a question that another bot can handle. Internally this probably maps to different tools or prompts, but as a user it is sometimes unclear which bot should receive which question. A single assistant with internal routing might feel more natural. The app could still have different capabilities behind the scenes, but the user would not need to choose between askbot, helpbot, newsbot, and weatherbot before asking a question.

Helpbot and product knowledge

Helpbot is meant to answer questions about the app, FAQ, and updates. It gave a sensible answer when I asked whether it could see the Wi-Fi signal strength on my phone. It said it could not access the phone’s hardware information, which is expected.

Helpbot Limitation

However, when I asked for the latest version of the app and how to update, helpbot did not know the exact version number or release date. My installed iOS app showed version 2.16 in the settings screen, but helpbot said it did not have the exact version number in its records and that its latest information was from May 8, 2026.

Bot not knowing the app version

For a general assistant this would be understandable, but for an app-specific helpbot it feels slightly incomplete. Since the bot is already inside the app, it would be useful if it could answer basic product questions such as the installed version, current platform, update instructions, and available settings.

Message deletion

I noticed that I could delete my own question, but I did not see an option to delete the bot’s answer. This leaves the conversation in a strange state because half of an exchange can be removed while the other half remains.

With Options to Delete Questions

This matters more in a voice app than in a normal text chat because voice messages may contain accidental recordings, personal details, or transcripts that were later corrected. Deleting an entire exchange would feel more predictable than deleting only one side of it.

Some wording choices

The label “new someone” appears when adding a new person.

New Someone Add Contact

Maybe this is intentional and part of the app’s playful tone, but it sounded unusual to me. “New contact” is more ordinary, but also clearer. The app already has a lot of personality in its visual design and voice interaction, so the core navigation labels do not necessarily need to be unconventional.

Email notifications

During testing, I noticed that the application sent me a new email every time the chatbot generated a reply. For human-to-human voice notes, email notifications make sense because email is part of the delivery mechanism. For bot conversations, it can become noisy quite quickly.

Email message received

If the user is actively talking to askbot, weatherbot, or newsbot inside the app, receiving an email for every single bot reply may be unnecessary. This seems like something that could be controlled by more detailed notification settings, especially if Talk Machine is used frequently.

Text on screen and spoken playback

The spoken playback does not always match the visible text exactly. The generated voice sometimes adds natural filler such as “uh”, “ah”, or “so”, and in some cases it slightly changes the wording. For example, a simple word like “dry” may become “it’s looking pretty dry”, while “you might want your umbrella” may become something closer to “you might want to grab your umbrella”.

I understand why this happens. The system is probably trying to make the spoken response sound more natural. In casual bot replies, that can be pleasant. However, it also means the text on screen and the spoken audio are not always the same message. For bot replies, this may not matter much. For human-to-human voice notes, the difference is more noticeable. If a message is being sent to another person, the transcript and spoken playback should ideally represent the same content.

Web audio playback

I also tested the web playback flow. While playing audio through the browser portal, I saw some CORS-related errors. This did not look like anything secretive or serious, and did not affect playback in my test.

Screenshot 2026-07-04 221518

Looking at the browser network tab, the audio stream appeared as an MPEG stream. A few years ago, I would probably have written a small script to decode the stream, validate the media header, and save the result into an audio file. With ChatGPT, that kind of conversion can now be done in a heartbeat. Some of the small utility programs I used to write by hand are now simply gone, or at least no longer worth writing from scratch.

ChatGPT decoded MPEG trivial

General impression

Talk Machine feels like an early product, but it also feels like a product with a clear idea behind it. The main strength is that it treats voice as the primary interface instead of treating it as an optional input method. The big talk button, waveform cards, voice setup, built-in bots, and email voice-note flow all support that direction.

The app is not just a chatbot with speech-to-text added. It is closer to a voice messaging system with AI features built into the conversation flow. The email integration is especially interesting because it lets the product communicate with people outside the app, which many messaging products struggle to do.

There are still rough edges. Some bot responsibilities overlap or feel too separated. Helpbot does not yet seem fully aware of the app itself. Email notifications can be noisy. Web audio playback needs more polish. The spoken voice sometimes differs from the text shown on screen. A few labels and message-management options also feel unfinished.

Even with those issues, the app was enjoyable to test. The onboarding was quick, speech recognition was better after refinement, Vietnamese news translation worked well, and the email-based voice note flow was genuinely fun. The product is still young, but the direction is promising.

0.00 avg. rating (0% score) - 0 votes
ToughDev

ToughDev

A tough developer who likes to work on just about anything, from software development to electronics, and share his knowledge with the rest of the world.

Leave a Reply

Your email address will not be published. Required fields are marked *

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <strike> <strong>