How Welocalize Turned a Million Manual Tasks Into an AI Agent Workflow artwork
Eventual Consistency | Your Reality Check on What's Actually Happening in Data

How Welocalize Turned a Million Manual Tasks Into an AI Agent Workflow

  • S2E13
  • 51:29
  • July 29th 2026

What happens when an AI agent's mistake can cost you the client, not just the demo.

You've proven your AI agent works in a sandbox. Getting the business to trust it in production, with real customers watching, is a different problem entirely.

Matthew Sekac is Head of Data & Analytics at Welocalize, where he has spent the past 18 months building an AI agent framework to handle the project management complexity behind thousands of translation and localisation projects. He joined Welocalize in 2020 and has over a decade of senior sales strategy experience in IP and life sciences translation before moving into data leadership.

Matt walks through exactly how his team proved an AI agent framework could handle over a million manual tasks a year without a mountain of rules-based technical debt. You'll get a concrete method for testing agent reliability against human benchmarks before you ever touch production and a way of thinking about human-in-the-loop verification that scales instead of adding new work.

This one covers ethnographic research into how project managers actually work, shadow testing agents against human output, and the evals-based thinking needed to get leadership buy-in for AI at scale. It's built for data and analytics leaders trying to move AI agents from pilot to production, not for anyone looking for a quick AI hype fix.

Key Takeaways

- Welocalize ran a 30-day shadow test where an AI agent mirrored every human decision on live tasks before a single customer saw it, and that proof, not confidence, is what got leadership to say yes.

- The team found their "quality check" agent was inversely correlated with actual accuracy: the more confident it was that a human wouldn't override it, the more likely a human did. The reason why says a lot about where reasoning agents still fail.

- Nobody asks if a human process is error-free before automating it, but everyone expects zero errors from AI. Matt argues that a mismatch is the real reason AI trust conversations go sideways.

- Before writing a single rule, the team spent weeks just watching project managers work, uncovering a 56-page SOP that turned into a 70-row spreadsheet for one customer alone.

Chapter Markers

00:00 Introduction: capability and trust as one problem

01:07 The agentic AI initiative at Welocalize

03:07 Why rules-based automation hit a technical debt wall

06:45 Decomposing project complexity: common cause vs special cause

10:04 The ethnographic study: watching before automating

12:32 Building the shadow test: 30 days, human vs agent

16:04 Real hiccups: parsing failures and UI friction

19:00 Where human-in-the-loop verification goes next

22:52 Benchmarking against human error, not perfection

26:07 Trust, blame, and why systems feel different to people

29:22 Gathering evidence leadership will actually believe

38:42 Building a repeatable agentic engineering practice

40:50 How AI is changing who can build automation

49:00 Evals-based development as a precondition, not an afterthought

Useful Links & Resources

- Matthew Sekac on LinkedIn: https://www.linkedin.com/in/matthew-sekac-8884894

- Welocalize: https://www.welocalize.com

Connect With the Show

- Host LinkedIn: https://www.linkedin.com/in/b-ross-katz/

- Host on X: https://x.com/brosskatz

- CorrDyn on LinkedIn: https://www.linkedin.com/company/corrdyn/

- Website: https://corrdyn.com

If you're wrestling with the same problem, proving an AI system's reliability before anyone will trust it with production work, we want to hear how you're approaching it. Drop a comment with what your organisation's baseline for "good enough" actually is.

If you want to talk about your data challenges, or you think we got something wrong, find us at corrdyn.com.

Eventual Consistency | Your Reality Check on What's Actually Happening in Data

The data leader's fortnightly reality check. No hype. No hot takes for engagement. Just honest conversation about what's actually happening in data and what it means for the work you're doing.

Every two weeks, we pick the stories dominating your feed, the acquisitions, product launches, frameworks, and controversies and discuss them the way you would with your team: critically, honestly, and with one question in mind: "What does this actually mean for my world?"

We're not here to sell you courses, predict the future, or tell you the sky is falling. We're here to cut through vendor claims that everything is "revolutionising" something, LinkedIn posts oscillating between doom and humble-brags, and tech journalism that treats every product launch like it's world-changing.

This is for VPs of Data, Analytics Directors, Data Engineering Managers, and senior practitioners who need to stay informed but don't have time to wade through whitepapers and noise. People making real decisions: Should we migrate to that warehouse? Is this ML use case worth it, or just shiny object syndrome? Why is everyone talking about this framework when it doesn't solve our actual problem?

In 20 minutes, you'll know what's worth your attention and what you can safely ignore. You'll get the perspective to make better decisions, ask vendors better questions, and avoid getting swept up in whatever trend is dominating feeds this week.

You'll hear from practitioners and consultants who've been in the room when these decisions go right and when they go spectacularly wrong. We know what the press release says. We also know what actually happens six months later.

Because in data, like in distributed systems, consistency is hard. But eventually, reality catches up with the hype.