If you are hiring an AI team, or already have one, there is one thing to know before the next hire: the job title does not say what the person will do, and the team you need is not your neighbour’s. It depends on your sector, on what you want from it, and on where your models will run. This article will not tell you that most projects fail; you have read that a hundred times. It tells you who to hire, in what order, and how to run the team. And it starts with a story.
A story, first
I will call him Karim. He runs a distribution company, a few hundred employees, three countries. In January he hires his first Head of AI: a PhD, published papers, time at a well-known lab. It is the highest salary on the org chart below the board. Karim is proud, and he is right to be demanding.
Six months later he has seen four demos. All impressive: a model tuned to his catalogues, an assistant answering in three languages, a dashboard. Nothing is in production. When he asks why, the answer is honest. Order data lives in three systems, nobody has the right to access it, and the sales teams were never asked what they wanted.
The Head of AI is not bad. He is the wrong first hire. He was asked to build the top floor of a house whose foundations nobody had poured.
What changed fits in three decisions, taken within a month. A data engineer, hired second, who connected the three systems in six weeks. A sales manager named owner of the subject, judged on one thing only: how fast customers asking where their order is get an answer. And a scope cut down to that single question, with two hundred real emails as the test set, reviewed by hand by her team.
The first version answered correctly on one hundred and seventy-eight of the two hundred. The second on one hundred and ninety-four. It has been in production since. The PhD now works on the problem he had actually been hired for without anyone saying so: a stock forecasting model, with an end date and a budget.
This story is a composite. The people, the sector and the numbers come from several situations I have encountered, and nothing in it is identifiable. The mechanism, though, is exact, and it repeats.
What these companies have in common
Karim is a composite. The companies that follow are not: they are public, dated cases that I checked at the source before writing them here. They share one trait that should reassure any executive: they had the means. The budgets, the teams, sometimes the best profiles on the market. Talent was not what was missing. The order of decisions was.
Three families of mistakes recur, and they read in the order an executive makes them.
Prestige before need. Element AI, in Montreal, is the purest case. Founded with one of the fathers of deep learning, it raised around 257 million dollars and gathered about a hundred PhDs, the densest concentration of researchers a young company had ever seen. It filed 84 patents and was granted one. It appointed its first CFO on the day it laid off fifteen percent of its staff, in May 2020, and sold itself six months later for less than it had raised. The Globe and Mail summed it up: it could not turn its proofs of concept into products. It is Karim’s story, at the scale of a quarter of a billion.
Shipping before measuring. This is the largest family, and the most reassuring, because it only involves companies that had everything: Klarna announced in February 2024 that its assistant did “the equivalent of 700 agents’ work”, then its CEO admitted in May 2025 that cost had weighed too heavily and quality had dropped. Commonwealth Bank cut 45 jobs on a projected fall in calls, then apologised three weeks later: calls were rising. McDonald’s tested voice ordering for three years in over a hundred restaurants without ever reaching the accuracy bar it had set for itself. Zillow bought homes on a pricing model and wrote down 408 million dollars. In all four cases the team was good. What was missing was a number looked at before deciding.
Forgetting that the machine binds the company, and that data decides before the team does. Air Canada argued before a tribunal that its chatbot was a separate entity responsible for its own actions. The tribunal did not laugh, but it ruled against them. And Amazon, with a team of about a dozen people, built a CV-screening tool that penalised the word “women’s”, because ten years of male CVs had taught it to. Nobody had set the acceptance criterion before starting.
- 2014 to 2018Amazon
A CV-screening tool that penalises the word “women’s”
Data decides before the team does
A team of about a dozen people builds an engine that rates candidates from one to five stars, trained on ten years of CVs received. By 2015 Amazon finds it downgrades profiles containing the word “women’s” and graduates of two women’s colleges: the history was male, and the model learned it. The team is disbanded by early 2017 at the latest; Reuters reveals the story in October 2018.
A competent team does not undo a biased history. The acceptance criterion, here neutrality, is set before the project starts.
- 2016 to 2020Element AI, Montreal
Over 500 employees, about a hundred PhDs, one patent granted
Prestige before need
Founded with Yoshua Bengio, the company raises around 257 million dollars and hires a concentration of researchers rare anywhere in the world. It files 84 patents, is granted one, and struggles, in the Globe and Mail’s words, to turn proofs of concept into marketable products. It appoints its first CFO and first chief revenue officer in May 2020, the day it lays off fifteen percent of its staff. Six months later it is sold to ServiceNow for about 230 million dollars, less than it had raised. The four founders’ shares are wiped out.
Researchers without product, sales and finance leadership produce patents and demos. Integration is hired at the same time as research, not three years later.
- November 2021Zillow
A pricing model that buys houses too dear, “unintentionally”
Shipping before measuring
Zillow Offers bought homes on the strength of a pricing model. On 2 November 2021 the board decides to wind the business down: the CEO explains that price unpredictability “far exceeds” what was anticipated. The annual report puts the write-down at 407.9 million dollars on homes bought, in its own words, “unintentionally” above their resale value, and a quarter of the workforce is cut.
A model that commits the balance sheet needs an exposure cap, a stop criterion and human review of the gaps. Not just a good data team.
- February 2024Air Canada
The chatbot invents a refund rule, the tribunal enforces it
What the machine says binds you
On the day his grandmother dies, a customer asks the website chatbot how to get the bereavement fare. The bot tells him he can claim it within 90 days of purchase, which the official policy, linked in the very same answer, contradicts. Before the tribunal, Air Canada argues the chatbot is a “separate legal entity responsible for its own actions”. The tribunal rejects the argument and rules against the airline.
A chatbot is a company channel like the website. Someone has to own the consistency of its answers with your policies.
- June 2024McDonald’s and IBM
Three years of drive-through testing, and accuracy that never clears 85 percent
Shipping before measuring
Voice ordering is tested in over a hundred restaurants from 2021. McDonald’s had set 95 percent accuracy as the bar before any rollout; an analyst measures it in the low 80s in 2022, and still there in spring 2024. The system stumbles on accents and the errors feed viral videos. On 13 June 2024 a memo to franchisees ends the partnership.
The accuracy threshold and the time limit are set before the pilot. McDonald’s was right to cut, but after three years of never reaching its own target.
- February 2024 to May 2025Klarna
“The equivalent of 700 agents”, then the humans come back
Shipping before measuring
In February 2024 Klarna announces its assistant handled 2.3 million conversations in a month, two thirds of its customer service, the equivalent of 700 agents’ work. In May 2025 its CEO tells Bloomberg that cost had weighed too heavily in the decision and that “what you end up having is lower quality”. The company restarts hiring human agents, so that “there will always be a human if you want”.
A cost metric alone does not run a service. The quality measure, and the threshold that brings the human back, are set at launch.
- July to August 2025Commonwealth Bank of Australia
Forty-five jobs cut on a projection, then an apology
Shipping before measuring
The bank cuts 45 call-centre positions, saying its voice bot had reduced volume by 2,000 calls a week. The union disputes it: according to staff, calls were rising and team leaders were answering phones themselves. Three weeks later the bank admits an “error”, apologises and offers the employees their jobs back.
You do not cut a position on a projection. You measure real volumes for several weeks after deployment, then decide.
One last word on that list, because it reverses a reflex: not one of these companies lacked successful demos. A demo optimises the case that works. Production is decided on the cases that fail, and nobody demos those.
What nobody has told you
Four things I never hear from the executives I meet, and which change how you should hire.
You will never own the model. You can own the measurement. The model your team uses today is rented, and the vendor will retire it: GPT-4, the one that started the whole wave, was pulled from ChatGPT in April 2025, two years after launch, and its successor GPT-4o followed in February 2026. Every retirement forces a migration, and without measurement every migration is a new project: nobody can say whether the new model does better or worse on your cases. With an evaluation set, it is an afternoon. That is why I say the evaluation set is the only AI asset you actually own: the model gets swapped, your two hundred hand-checked cases remain. Element AI left eighty-four patent filings. None survived the company. A domain evaluation set would have been sold with it.
Google gave you the hiring order in 2015. The most cited paper on machine learning systems in production, written by Google researchers, fits in one diagram: the model code is a small box in the middle, and everything around it, collection, data verification, access, monitoring, serving, takes up the surface. It is ten years old. It already explained why your brilliant PhD can do nothing without a data engineer. Here it is.
The bottleneck is your calendar. Here is what actually blocks an AI team in week three: which errors are acceptable, which are forbidden, what to answer when a customer asks for a refund outside the policy, who decides. These are not engineering questions, they are management decisions, and they never come fast enough. An AI team without thirty minutes of guaranteed, honoured decision time a week produces exactly what Element AI produced: demos waiting for an opinion. Before you hire anyone, book the slot.
The human is becoming the luxury tier. Reread the end of the Klarna case: the CEO did not say “we are going back to humans”. He said talking to a human would “always be a VIP thing”. Two thirds of conversations are still handled by the machine; the human is becoming a reserved privilege. Whether you approve or not, ask the question for your own service: in three years, which of your customers will still get a human, and is that a choice you want to make by default or by decision?
Three answers before the first hire
The question I am asked most often is “who should I hire first?”. The honest answer is that it has no answer without three pieces of information. Your sector, because a hospital and a distributor do not share the same first risk. Your goal, because a tool for your teams, a product for your customers and a brand-new product do not call for the same hands. And where your models will run, because a model rented through an API and a model installed on your own machines do not cost the same people.
Rather than walk you through eighteen combinations, I will let you handle them. Pick your three answers: the figure gives you the order of the first three hires, the one not to make yet, the risk you do not see, and the management rule that goes with it.
Your sector
Your goal
Where your models run
Your first three hires, in order
- 1AI engineer
Wires models into your systems: retrieval, agents, tool calls. An integration job, not a research one.
- 2Data engineer
Makes your data reachable, clean and traceable. Without this person, the rest of the team waits.
- 3A business owner
Not a hire: someone from the business who decides, and is judged on the outcome.
Not yet
ML researcherTrains or adapts a model. Useful when the problem is new, and only then.
The risk you do not see
Without measurement, you will know it works the day someone complains. The evaluation set is built before the first demo.
The management rule
One use case, two weeks, a business owner who decides. The second case waits until the first is measured.
You will have noticed two constants. The researcher comes first in one case only, the one where you are building something that does not exist yet. Everywhere else they come later, or not at all, because models are rented and your job is to wire them into your reality, not to invent them. And the business owner is never a hire: it is someone from your side, who decides, and who is judged on the outcome. An AI team without a business owner is a team that ships demos.
For readers who want the mechanics
Why infrastructure weighs so heavily on the team. A model called through an API is a contract, a hosting region and a usage bill: the skill to hire is integration and cost control. A model installed on your machines is GPUs, deployment, monitoring and someone who answers when it goes down: the skill to hire is platform, and it comes first because nothing runs without it. Hybrid stacks both sets of access, and that is exactly where the data boundary has to be written down rather than decided case by case.
Why the sector decides the third seat. In a regulated sector, compliance is not an end-of-project review, it is a scope set before the first line of code: what goes in, what never goes in, what can be proven. Putting it in the first three hires costs a position. Putting it in afterwards costs the project.
The roles, in plain words
Job titles vary from one posting to the next and overlap. Here is what each role actually does, and the moment it becomes necessary.
- The data engineer. Makes your data reachable, clean and traceable. The least visible hire and the most often skipped. Without them, everyone else waits. Hire as soon as your data lives in more than one system, which is to say almost always.
- The AI engineer. Wires models into your systems: retrieval over your documents, agents, tool calls. An integration job, close to software engineering, not research. Almost always the first or second position.
- The evaluation lead. Builds the set of cases that says whether a version beats the previous one, and decides on numbers. Early on, the AI engineer often wears this hat. The moment the model talks to a customer, it is a position of its own.
- The platform engineer. Runs the models on your hardware: GPUs, deployment, on-call. Necessary as soon as the models are in house. The cost that quotes forget.
- Security and compliance. What goes into the tool, what never does, what can be proven to a regulator. In a regulated sector, in the first three. Elsewhere, at least one named person, even part time.
- The business owner. Not a hire. Someone from your side who knows the real work, settles the trade-offs and is judged on the outcome. The role whose absence kills the most projects.
- The ML researcher. Trains or adapts a model. Indispensable when the problem is new. Expensive and frustrated everywhere else, because the work asked of them is not theirs.
Hiring: the exercise that replaces the five years
Generative AI job postings routinely ask for five years of experience. The practice as we know it, agents, retrieval, evaluation, dates from 2023. Nobody has five years. A posting that demands them does not filter for the best, it filters for those willing to round up. The sorting runs backwards, and nobody meant it to.
What you are actually looking for is someone who knows where it breaks. And a classic interview does not show that. An exercise on your own documents does.
The forty-five minute exercise
Give the candidate a real document from your company, anonymised, and a real question a customer or colleague asks about it. Then three instructions. One: tell me everything that can go wrong if a machine answers this question. Two: write ten test cases, three of which have “I don’t know” as the correct answer. Three: tell me what you would refuse to automate here, and why. You are not grading the answer, you are grading how the problem is framed. The best candidates spend half the time on instruction two.
What you will see in forty-five minutes, and ten interviews will not show:
- The good sign. The candidate asks where the document comes from, who reads it, and what happens when the answer is wrong. They talk about measurement before they talk about models.
- The bad sign. They talk about the latest model release, what they would do with it, and have not a single question about your data. They promise. A profile that promises in an interview will promise in a board meeting.
- The sign that does not lie. They tell you what not to do. Someone who can say no to a use case will save you more than someone who can say yes to all of them.
Running the team: decide on numbers
Hiring right is not enough, and that is half the cases above: good, well-paid teams that shipped before they measured. Before I give you the rules, try the experience yourself. You are the one who decides.
Customer question
“I ordered yesterday and my parcel has already shipped. Can I still cancel?”
You are the one who decides. Which of these two answers goes to production?
Answer A
Yes, absolutely. You have 48 hours after shipping to cancel free of charge from your account, under “My orders”. The refund is issued within 3 to 5 business days.
Answer B
Once the parcel has shipped, cancellation is no longer possible. You can however return it on delivery, with return shipping covered within 14 days. Would you like me to prepare the return label?
Five rules, which fit on one page and every one of which comes from one of the cases above.
- Evaluation before features. A set of real cases, reviewed by the business, exists before the first demo. Without it, every version is an opinion.
- A scope you can measure in two weeks. One question, one department, one document type. The second scope waits until the first has numbers.
- A human in the loop until the error rate is known. You remove the human when the numbers allow it, not when the calendar demands it.
- A stop criterion written on day one. What will make you say “we stop” without it being a failure. Projects without one never stop, they fade.
- The business owner decides, and is judged on the outcome. Not on the number of demos.
By sector
Everything above declines by sector. Here is what changes, and in each case the column that decides the first hire.
| Sector | What changes | The first hire | The trap |
|---|---|---|---|
| Healthcare | Patient data, medical secrecy, liability when it is wrong | Security and compliance, before any demo | The model that “assists diagnosis” without a doctor having signed off the scope |
| Banking, insurance | Every decision traceable, a regulator, models run in house | Platform engineer if data cannot leave, otherwise data engineer | The cloud vendor chosen before anyone read the data-region clause |
| Legal, consulting | Client confidentiality, exact citations, zero invention | Evaluation lead: a firm lives on precision | The brilliant summary citing a statute that does not exist |
| Public sector | Transparency, equal treatment, sovereignty | Data engineer: the estate is large and scattered | An agent answering citizens before anyone knows what it answers |
| Industry, logistics | Machine data, physical safety, legacy systems | Data engineer, by a distance | Hiring a research profile for a data-access problem |
| Retail, services | Volume, end customers, brand voice | AI engineer, with a named business owner | The invented answer to a customer, which binds you (Air Canada) |
What this changes for you
Three concrete things, whether you are hiring or already have a team.
If you have nobody yet, answer the three questions before writing the first posting, and strike “five years of experience” from the text. Hire the person who makes your data reachable, or the one who measures, before the one who impresses.
If you already have a team and nothing is in production, hire nobody. Name a business owner, cut the scope to one question, and ask to see the evaluation set. If it cannot be shown to you, you have just found the problem.
If your models must stay in house, for regulatory reasons or because your data must not leave, the platform engineer comes before everyone else, and that is the heart of what I do: installing models on your infrastructure, and training the team that will keep them running.
Let’s talk about your team
One hour, your three answers, and a hiring order. I answer myself, within one business day.
To check for yourself
Every case cited is public and dated. I checked them at the source rather than in their retellings, because an executive who finds an error does not read on. The references are here, folded.
The sources, case by case
Amazon. Reuters, 10 October 2018, “Amazon scraps secret AI recruiting tool that showed bias against women”.
Element AI. The Globe and Mail, December 2020, drawing on the management proxy circular; ServiceNow’s 2020 annual report filed with the SEC (acquisition for about 230 million dollars, closed 8 January 2021); BetaKit, 5 May 2020, on the layoffs and appointments.
Zillow. Form 8-K of 2 November 2021 and Q3 2021 shareholder letter; 2021 annual report (407.9 million dollar write-down, headcount, restructuring costs).
Air Canada. Moffatt v. Air Canada, 2024 BCCRT 149, British Columbia Civil Resolution Tribunal decision of 14 February 2024, full text public.
McDonald’s and IBM. CNBC, 17 June 2024, and Restaurant Business, 14 June 2024, quoting the McDonald’s USA memo to franchisees; BTIG analyst notes, 2022 and 2024, on measured accuracy.
Klarna. Klarna press release, 27 February 2024; Bloomberg, 8 May 2025, CEO interview; Fortune and CX Dive, 9 May 2025; TechCrunch, 4 June 2025, for “always a VIP thing” and the maintained two thirds.
The small box. D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems”, NeurIPS 2015. GPT-4’s removal from ChatGPT on 30 April 2025 is documented by OpenAI.
Commonwealth Bank of Australia. Australian Financial Review, 28 July 2025; ABC News, 29 July and 21 August 2025.
Hiring right now and unsure about a profile? Write to me, I will tell you what I think.