What If AI Sat in the Middle of Every Conversation?
I remember the exact moment my ex-co-founder sent me an 11-minute-36-second WhatsApp audio. And I was thinking: holy fuck, now I have to listen to this.
I am a heavy WhatsApp user, and before AI, I really hated when people sent voice messages. First reason: WhatsApp didn’t have 1x, 2x playback, and my brain thinks pretty fast. Second reason: no transcript. So an 11-minute audio was just… 11 minutes of my life, gone, at someone else’s pace.
Voice went from a burden to a superpower
Then WhatsApp introduced 1x, 2x speed. And, to be honest, it really helped — for more than one reason. First, obviously, speed. But second, even the tone. When you listen to an audio, you can really understand what the other person is thinking. It’s deeper than writing. And that’s the thing people love and hate about it, for the same exact reason.
I remember listening to one where the tone changed, and I could tell the moment wasn’t good with my ex-co-founder. That would never have come through in text.
A WhatsApp audio message contains way more shades than the 50 shades of that well-known movie. Now, whenever I need detail, I explicitly ask people to either create a “room” — so I get the audio and an immediate transcript — or send me a proper voice note. Room is better, honestly, because you get both. But a raw WhatsApp audio still gives me more shape than text ever will.
The bigger question: what about a whole platform built on this?
So what about a new social platform, or a working tool — Slack-ish — built voice-first from the ground up?
Here’s why this is connected to my own story. If I had to define myself in one tagline, I’d say: Alessio Ragni, trying to understand humans since 1985 and still doing mistakes. For me it’s genuinely difficult to understand humans, because they’re not logical. We are not logical. That’s the reality. Sometimes people act almost like animals — but that’s another story.
The reality is that most people don’t care about each other in how they communicate. Maybe they’re aggressive because they want to be aggressive, without caring about how it lands. Maybe they write without checking grammar, or without checking anything at all.
Let’s take a remote environment. I’m not an English native speaker, but I used to work with American companies, and in Slack I used to tweak every single message I sent. And still, sometimes they’d write slang I didn’t understand, and I’d be stuck.
This is also connected to a failure I had in the past: Alfred
Alfred was a startup I built with a similar expectation. The idea was: you write something, and Alfred rephrases it before you even hit send.
Let me pitch it to you the way I pitched it back then. Imagine you’re about to send a text to a colleague, and instead it just replies: “Hey mate, yes, I agree with you,” or something like that — cleaned up, before it goes out.
Now push that idea forward: what if the AI sits in the middle, not before you send, but between you and the receiver?
The other side never sees your raw audio. They don’t even see a transcript. They see — well, maybe they can reply to the robot version, but that’s another story — a translated version. Passed through AI, passed into the format they prefer to receive information in. The AI sits in the middle. And that’s the important part.
I don’t care if the AI is “judging” or “advising” — I don’t care about the label. What matters is that it acts the same way for both sides. It’s a kind of arbiter between two parties. It will simplify or adjust to the style of the reader. But for the person sending the message, the reality doesn’t change: they can send however they want. Raw. Unfiltered. However it comes out.
Let’s take the worst case ever. You’re sending me something and you want to say, “Hey, you are bullshit!” I mean — that’s not true, and I’d probably reply with something just as blunt. The AI translates it into: “This is inaccurate.” Which is not totally polite either, okay — but it’s way more efficient than calling someone stupid.
And yes, that’s the worst example I could give you. But imagine if the AI can actually read the intent, understand what you’re trying to say, maybe even ask you a clarifying question — and then send the right message. Sure, we’ll burn a lot of tokens doing it. But look at the power of that.
On one end, humans get to fully express themselves. On the other end, AI gives humans 10x — it doesn’t replace them.
There’s a downside, obviously
Humans might stop learning how to improve. They might stay dragged along, hiding behind a keyboard — and then when they actually have to sit face to face with another human, they won’t be able to do it.
Can I be honest? That’s already happening. Right now, without any of this.
And once again — technology, to me, is agnostic. It’s not good or bad in itself. It depends entirely on how you use it. I’d even use something like this when I need to discuss something with my girlfriend — though in that case, she’d probably rather just listen to my actual voice, which by the way is so sexy. But that’s another story.
Why I keep talking with AI every single day
I’m a product engineer now — I really don’t want to be called a coder, to be honest. And as a product engineer, we’re not writing code anymore. We’re discussing. Back and forth, all day.
So imagine the power of an AI that doesn’t just sit there transcribing — one that actually sits in the middle, between every sender and every receiver, translating not the words but the intent.
I don’t know if I’d build it. But I keep coming back to it.
I love feedback.
Did this help? One tap — that's it.
Thanks 🙏
Want to add anything — or get in touch? Totally optional.
Got it — thank you. 🙏