Hey team,
In a couple of hours, we head out on a flight to Dubrovnik, rent a car to drive to Kotor, and spend a few days in the Balkans for a quick change of scenery.
The most fun part of this trip is that this is the first trip I have fully entrusted to the bots to book entirely for me… let’s see how it goes 🙃
Instinct found me a great points flight after checking for United saver redemptions daily for a few weeks.
Muse used my Chase Ultimate Rewards points to book a “points drop” hotel in Dubrovnik; the cash price was $1,500 and we used 64k points plus the $250 EDIT credit from Chase Sapphire Reserve.
Muse also went back and forth negotiating with countless Airbnbs until we got a great deal in a top 1% listing in Kotor old town, saving more than 20%.
P.S. First few folks to grab this Muse invite, add code CRX40I in settings for 1B tokens. Instinct invite here too. You’re welcome.
Handing a bot my vacation is easy when the job is clear, and the bot can actually finish it. CX gets messier when the bot writes a beautiful answer, can’t take the action, and still gets counted as a win.
A customer asks for a refund, gets a polished explanation of the refund policy, and comes back two days later to find a human asking for the same order number they already gave twice.
Somewhere, that first conversation is sitting in a report as an AI success.
Let’s get into it.

This week’s newsie is brought to you by Kim.cc.
I’ve been thinking a lot about the gap between *having AI in your support stack* and actually having AI that gets better over time.
NovaaLabs is a pretty good example.
They’re an 8-figure ecommerce brand and had been using Gorgias AI for support for over a year.
And the frustrating part wasn’t that the AI made mistakes. Obviously it will.
*It was that it kept making the same mistakes.*
Their team would catch something, correct it, and then run into a version of the same problem again.
Which raises a pretty important question:
If your team constantly has to watch, correct, and retrain the AI… how much work did you actually automate?
Ahead of BFCM, NovaaLabs moved to Kim.cc.
Within the first month, *45%+ of their support tickets were being handled through Kim, with less than a 3% re-open rate.*
And maybe even more importantly, their internal team did almost none of the heavy lifting to get there.
Kim had CX, AI training, and implementation experts handling the setup, training, QA, and ongoing improvement of the system.
That distinction is interesting to me.
Most AI support products give your team a tool and then leave you responsible for making that tool successful.
Kim is an *AI-native service*, which means the AI handles the tickets. Their experts, also known as "Sentinels", make sure it’s learning, improving, and not repeating the same mistakes over and over again.
So NovaaLabs went from spending time managing the AI to having *45%+ of their tickets handled within a month, <3% reopens, and close to zero implementation lift from their own team.*
Especially heading into BFCM, I think that’s the metric more teams should be looking at.
“How many did it resolve well enough that the customer didn’t have to come back?”
If you’re thinking about AI support before peak season, Kim.cc is worth a look.
Please let me fix the easy thing myself
I’m comfortable letting an agent spend weeks checking flight availability because I have absolutely no interest in doing that myself. The dates and constraints are clear, and there’s a pretty obvious definition of success.
If we land and the hotel has no record of our booking, my enthusiasm for experimentation will depend heavily on whether someone can get us a room.
The same customer can love automation for one part of an experience and need a person five minutes later.
That makes considerably more sense to me than declaring a brand “human-first” or “AI-first” and forcing every situation to fit the announcement.
The self-service experiences at Quince, Wayfair, Target, and Amazon are my favorite examples. For eligible issues, you can go into an order and work through the available resolution options without even opening a conversation with a human.
For a straightforward exchange, I’d take that over explaining the same thing four times to someone who then needs to “check their system.” Let me pick the right size and get the return instructions, and I’m perfectly happy.
This is where anyone starting with AI should slow down for a second, because some of your biggest wins are boring self-service fixes.
Before you build a bot that can discuss your return policy in eleven different tones, make sure customers can actually start a return.
AI earns its keep on the messier requests: someone doesn’t know which order contains the item, explains three problems in one message, or needs help figuring out which option fits.
But it needs the data and permissions to finish. Otherwise, you’ve built something that can confidently describe a refund while being entirely unable to issue one.
The customer still has work to do, however nice the explanation was.
What exactly are we calling resolved?
If you reward a system for keeping conversations away from humans, as we lovingly call “deflection” in this industry, you need to know what happens to the people who never reach one.
Some got exactly what they needed. Others may have given up, tried another channel, or decided they’d deal with it tomorrow.
If your reporting treats a conversation ending without a human as a success, those outcomes can look remarkably similar.
The bot recites the policy, the customer closes the chat, and then emails support the next morning. Unless you connect those interactions, you end up celebrating the first one while paying someone to handle the second.
I don’t think the people building these systems necessarily have bad intentions. They’re often working toward the target they were given.
But I would want “resolved” to involve checking that the action happened and addressed the request. A refund explanation and a processed refund should not earn the same little green tick.
Recontact and reopen rates are useful checks, too, provided you can follow the issue across channels. Silence alone isn’t proof that somebody went away happy.
Read the rest of the message
“Where is my order?” could mean the customer wants a tracking link.
It could also mean the replacement for a missing order has now disappeared, the customer leaves tomorrow, and someone on your team already promised this shipment would arrive in time.
The tag or subject line might be identical. Sending both customers the same tracking response would be a fairly bold choice.
Before the system keeps going, I’d want it checking:
Does it have reliable information? If the order record and carrier update disagree, what needs investigating?
Can it take an authorized action that solves the request? Explaining the policy doesn’t count as completing the action.
Has this already failed? A customer returning with the same unresolved problem should not be sent through the same sequence by default.
Would a mistake be hard to reverse or carry meaningful risk? Those cases need stronger checks and, where appropriate, a qualified person.
Has the customer asked for a human? There should be an obvious way to reach one without guessing the phrase that unlocks the door.
Frustration alone shouldn’t decide the route. An annoyed customer might be thrilled with an immediate automated replacement. A calm customer might be asking for something the system has no business approving.
I’d also be careful about creating a setup where becoming more upset is the fastest way to get useful help. Customers will learn that lesson.
A transfer to someone who can’t do anything is just a longer wait
I’ve spent enough time in CX to know how much a good person can do with a complicated situation.
They notice when the proposed fix doesn’t fit, ask the question nobody thought to ask, pull in another team, or make a judgment call within their authority.
Take a gift that arrived damaged the day before someone needed it. Another shipment arriving next week may be useless. A person can work through the alternatives with the customer instead of repeatedly offering the one option the system happens to support.
But that person needs access, authority, and a clear path to whoever can make the decisions they can’t.
If your rep can only repeat what the bot already said, the transfer has mostly changed who is delivering the disappointment.
AI adoption is a terrible excuse to make human help harder to reach. Lovely people are also not a good excuse for making customers wait through a slow, repetitive process. Either way, the customer ends up paying for how strongly you feel about your approach.
The customer should not have to train the next agent
This is the handoff I’d spend real time designing.
Before a person joins, they should see what the customer wants, what already happened, what was tried, and why the system stopped.
If the customer uploaded a photo, bring the photo. If someone promised delivery by Friday, surface that promise. If the bot couldn’t proceed because the records disagree, show the discrepancy.
Keep the original conversation available, too. The rep should be able to check the summary, and an AI suggestion should be clearly identified as a suggestion rather than something already approved.
“Customer is frustrated about delivery” leaves the next person redoing the investigation, now with a customer who has had additional time to become frustrated about delivery.
For a hypothetical damaged replacement, a useful opening would be:
“I can see this is the second damaged item and that you need it before Friday. I have the photo you sent. I’m checking whether we can get a replacement to you in time, and what we can offer if we can’t.”
That person has shown they read the conversation without promising something they haven’t checked.
The customer also needs to know who owns the next step and when they’ll hear back. If the team won’t be available until morning, tell them. “Connecting you now” should involve some actual connecting.
Pick one ugly ticket and follow it to the end
Start with one recurring issue your team already knows gets weird. Review real cases with the people who handle it. Find where the answer is clear, which details change the decision, and what a good escalation looks like.
Then test the ugly versions: out-of-stock replacements, conflicting records, previous failed attempts, customers who can’t use the standard return method, and tickets where the bot gave the right policy but the wrong resolution.
Check whether the action happened correctly, how much effort the customer spent, and whether they needed further help. When something escalates, look at whether the person had enough context and authority to move it forward.
Use what you find. If the same missing information keeps forcing a transfer, fix the information gap. If reps repeatedly wait for the same approval, look at their permissions. Those are decisions someone still needs to own after the software goes live.
I’ve entrusted the bots with quite a lot of this trip. If we arrive in Dubrovnik and the booking has gone sideways, I’m not going to care whether the fix comes from AI, a human, or a guy named Luka at the front desk.
I’m going to care whether someone understands what happened and can get us a room.
We can discuss the future of customer service after we’ve put the bags down.

That's it for this week!
Any topics you'd like to see me cover in the future?
Just shoot me a DM or an email!
Cheers,
Eli 💛




