All articles

Leading with AI

Hiring an AI team: three answers before the first hire

Who to hire, in what order, and how to run the team, depending on your sector, your goal and where your models run. With what well-known companies paid to learn it, and a story to start.

An org chart where three seats are lit

If you are hiring an AI team, or already have one, there is one thing to know before the next hire: the job title does not say what the person will do, and the team you need is not your neighbour’s. It depends on your sector, on what you want from it, and on where your models will run. This article will not tell you that most projects fail; you have read that a hundred times. It tells you who to hire, in what order, and how to run the team. And it starts with a story.

A story, first

A story

I will call him Karim. He runs a distribution company, a few hundred employees, three countries. In January he hires his first Head of AI: a PhD, published papers, time at a well-known lab. It is the highest salary on the org chart below the board. Karim is proud, and he is right to be demanding.

Six months later he has seen four demos. All impressive: a model tuned to his catalogues, an assistant answering in three languages, a dashboard. Nothing is in production. When he asks why, the answer is honest. Order data lives in three systems, nobody has the right to access it, and the sales teams were never asked what they wanted.

The Head of AI is not bad. He is the wrong first hire. He was asked to build the top floor of a house whose foundations nobody had poured.

What changed fits in three decisions, taken within a month. A data engineer, hired second, who connected the three systems in six weeks. A sales manager named owner of the subject, judged on one thing only: how fast customers asking where their order is get an answer. And a scope cut down to that single question, with two hundred real emails as the test set, reviewed by hand by her team.

The first version answered correctly on one hundred and seventy-eight of the two hundred. The second on one hundred and ninety-four. It has been in production since. The PhD now works on the problem he had actually been hired for without anyone saying so: a stock forecasting model, with an end date and a budget.

This story is a composite. The people, the sector and the numbers come from several situations I have encountered, and nothing in it is identifiable. The mechanism, though, is exact, and it repeats.

What these companies have in common

Karim is a composite. The companies that follow are not: they are public, dated cases that I checked at the source before writing them here. They share one trait that should reassure any executive: they had the means. The budgets, the teams, sometimes the best profiles on the market. Talent was not what was missing. The order of decisions was.

Three families of mistakes recur, and they read in the order an executive makes them.

Prestige before need. Element AI, in Montreal, is the purest case. Founded with one of the fathers of deep learning, it raised around 257 million dollars and gathered about a hundred PhDs, the densest concentration of researchers a young company had ever seen. It filed 84 patents and was granted one. It appointed its first CFO on the day it laid off fifteen percent of its staff, in May 2020, and sold itself six months later for less than it had raised. The Globe and Mail summed it up: it could not turn its proofs of concept into products. It is Karim’s story, at the scale of a quarter of a billion.

Shipping before measuring. This is the largest family, and the most reassuring, because it only involves companies that had everything: Klarna announced in February 2024 that its assistant did “the equivalent of 700 agents’ work”, then its CEO admitted in May 2025 that cost had weighed too heavily and quality had dropped. Commonwealth Bank cut 45 jobs on a projected fall in calls, then apologised three weeks later: calls were rising. McDonald’s tested voice ordering for three years in over a hundred restaurants without ever reaching the accuracy bar it had set for itself. Zillow bought homes on a pricing model and wrote down 408 million dollars. In all four cases the team was good. What was missing was a number looked at before deciding.

Forgetting that the machine binds the company, and that data decides before the team does. Air Canada argued before a tribunal that its chatbot was a separate entity responsible for its own actions. The tribunal did not laugh, but it ruled against them. And Amazon, with a team of about a dozen people, built a CV-screening tool that penalised the word “women’s”, because ten years of male CVs had taught it to. Nobody had set the acceptance criterion before starting.

  1. 2014 to 2018Amazon

    A CV-screening tool that penalises the word “women’s”

    Data decides before the team does

    A team of about a dozen people builds an engine that rates candidates from one to five stars, trained on ten years of CVs received. By 2015 Amazon finds it downgrades profiles containing the word “women’s” and graduates of two women’s colleges: the history was male, and the model learned it. The team is disbanded by early 2017 at the latest; Reuters reveals the story in October 2018.

    A competent team does not undo a biased history. The acceptance criterion, here neutrality, is set before the project starts.

  2. 2016 to 2020Element AI, Montreal

    Over 500 employees, about a hundred PhDs, one patent granted

    Prestige before need

    Founded with Yoshua Bengio, the company raises around 257 million dollars and hires a concentration of researchers rare anywhere in the world. It files 84 patents, is granted one, and struggles, in the Globe and Mail’s words, to turn proofs of concept into marketable products. It appoints its first CFO and first chief revenue officer in May 2020, the day it lays off fifteen percent of its staff. Six months later it is sold to ServiceNow for about 230 million dollars, less than it had raised. The four founders’ shares are wiped out.

    Researchers without product, sales and finance leadership produce patents and demos. Integration is hired at the same time as research, not three years later.

  3. November 2021Zillow

    A pricing model that buys houses too dear, “unintentionally”

    Shipping before measuring

    Zillow Offers bought homes on the strength of a pricing model. On 2 November 2021 the board decides to wind the business down: the CEO explains that price unpredictability “far exceeds” what was anticipated. The annual report puts the write-down at 407.9 million dollars on homes bought, in its own words, “unintentionally” above their resale value, and a quarter of the workforce is cut.

    A model that commits the balance sheet needs an exposure cap, a stop criterion and human review of the gaps. Not just a good data team.

  4. February 2024Air Canada

    The chatbot invents a refund rule, the tribunal enforces it

    What the machine says binds you

    On the day his grandmother dies, a customer asks the website chatbot how to get the bereavement fare. The bot tells him he can claim it within 90 days of purchase, which the official policy, linked in the very same answer, contradicts. Before the tribunal, Air Canada argues the chatbot is a “separate legal entity responsible for its own actions”. The tribunal rejects the argument and rules against the airline.

    A chatbot is a company channel like the website. Someone has to own the consistency of its answers with your policies.

  5. June 2024McDonald’s and IBM

    Three years of drive-through testing, and accuracy that never clears 85 percent

    Shipping before measuring

    Voice ordering is tested in over a hundred restaurants from 2021. McDonald’s had set 95 percent accuracy as the bar before any rollout; an analyst measures it in the low 80s in 2022, and still there in spring 2024. The system stumbles on accents and the errors feed viral videos. On 13 June 2024 a memo to franchisees ends the partnership.

    The accuracy threshold and the time limit are set before the pilot. McDonald’s was right to cut, but after three years of never reaching its own target.

  6. February 2024 to May 2025Klarna

    “The equivalent of 700 agents”, then the humans come back

    Shipping before measuring

    In February 2024 Klarna announces its assistant handled 2.3 million conversations in a month, two thirds of its customer service, the equivalent of 700 agents’ work. In May 2025 its CEO tells Bloomberg that cost had weighed too heavily in the decision and that “what you end up having is lower quality”. The company restarts hiring human agents, so that “there will always be a human if you want”.

    A cost metric alone does not run a service. The quality measure, and the threshold that brings the human back, are set at launch.

  7. July to August 2025Commonwealth Bank of Australia

    Forty-five jobs cut on a projection, then an apology

    Shipping before measuring

    The bank cuts 45 call-centre positions, saying its voice bot had reduced volume by 2,000 calls a week. The union disputes it: according to staff, calls were rising and team leaders were answering phones themselves. Three weeks later the bank admits an “error”, apologises and offers the employees their jobs back.

    You do not cut a position on a projection. You measure real volumes for several weeks after deployment, then decide.

Figure 1. Seven public cases, checked at the source, from 2014 to 2025. The same mistakes recur: prestige before need, shipping before measuring, and forgetting that what the machine says binds the company. References are at the end of the article.

One last word on that list, because it reverses a reflex: not one of these companies lacked successful demos. A demo optimises the case that works. Production is decided on the cases that fail, and nobody demos those.

What nobody has told you

Four things I never hear from the executives I meet, and which change how you should hire.

You will never own the model. You can own the measurement. The model your team uses today is rented, and the vendor will retire it: GPT-4, the one that started the whole wave, was pulled from ChatGPT in April 2025, two years after launch, and its successor GPT-4o followed in February 2026. Every retirement forces a migration, and without measurement every migration is a new project: nobody can say whether the new model does better or worse on your cases. With an evaluation set, it is an afternoon. That is why I say the evaluation set is the only AI asset you actually own: the model gets swapped, your two hundred hand-checked cases remain. Element AI left eighty-four patent filings. None survived the company. A domain evaluation set would have been sold with it.

Google gave you the hiring order in 2015. The most cited paper on machine learning systems in production, written by Google researchers, fits in one diagram: the model code is a small box in the middle, and everything around it, collection, data verification, access, monitoring, serving, takes up the surface. It is ten years old. It already explained why your brilliant PhD can do nothing without a data engineer. Here it is.

Data collectionData verificationFeature extractionConfigurationModel codeMachine resourcesAnalysis toolsAccess managementMonitoringServing infrastructure
Figure 2. After the diagram in “Hidden Technical Debt in Machine Learning Systems”, Sculley et al., NeurIPS 2015, Google researchers. In a real system, the model code is the small box. Ten years on, most job postings still hire for the small box.

The bottleneck is your calendar. Here is what actually blocks an AI team in week three: which errors are acceptable, which are forbidden, what to answer when a customer asks for a refund outside the policy, who decides. These are not engineering questions, they are management decisions, and they never come fast enough. An AI team without thirty minutes of guaranteed, honoured decision time a week produces exactly what Element AI produced: demos waiting for an opinion. Before you hire anyone, book the slot.

The human is becoming the luxury tier. Reread the end of the Klarna case: the CEO did not say “we are going back to humans”. He said talking to a human would “always be a VIP thing”. Two thirds of conversations are still handled by the machine; the human is becoming a reserved privilege. Whether you approve or not, ask the question for your own service: in three years, which of your customers will still get a human, and is that a choice you want to make by default or by decision?

Three answers before the first hire

The question I am asked most often is “who should I hire first?”. The honest answer is that it has no answer without three pieces of information. Your sector, because a hospital and a distributor do not share the same first risk. Your goal, because a tool for your teams, a product for your customers and a brand-new product do not call for the same hands. And where your models will run, because a model rented through an API and a model installed on your own machines do not cost the same people.

Rather than walk you through eighteen combinations, I will let you handle them. Pick your three answers: the figure gives you the order of the first three hires, the one not to make yet, the risk you do not see, and the management rule that goes with it.

Your sector

Your goal

Where your models run

Your first three hires, in order

  1. 1
    AI engineer

    Wires models into your systems: retrieval, agents, tool calls. An integration job, not a research one.

  2. 2
    Data engineer

    Makes your data reachable, clean and traceable. Without this person, the rest of the team waits.

  3. 3
    A business owner

    Not a hire: someone from the business who decides, and is judged on the outcome.

Not yet

ML researcherTrains or adapts a model. Useful when the problem is new, and only then.

The risk you do not see

Without measurement, you will know it works the day someone complains. The evaluation set is built before the first demo.

The management rule

One use case, two weeks, a business owner who decides. The second case waits until the first is measured.

Figure 3. The team is not the same across sectors, goals and where the models run. Three answers, and the hiring order changes. The rules are the author’s, open to challenge line by line, and that is the point.

You will have noticed two constants. The researcher comes first in one case only, the one where you are building something that does not exist yet. Everywhere else they come later, or not at all, because models are rented and your job is to wire them into your reality, not to invent them. And the business owner is never a hire: it is someone from your side, who decides, and who is judged on the outcome. An AI team without a business owner is a team that ships demos.

For readers who want the mechanics

Why infrastructure weighs so heavily on the team. A model called through an API is a contract, a hosting region and a usage bill: the skill to hire is integration and cost control. A model installed on your machines is GPUs, deployment, monitoring and someone who answers when it goes down: the skill to hire is platform, and it comes first because nothing runs without it. Hybrid stacks both sets of access, and that is exactly where the data boundary has to be written down rather than decided case by case.

Why the sector decides the third seat. In a regulated sector, compliance is not an end-of-project review, it is a scope set before the first line of code: what goes in, what never goes in, what can be proven. Putting it in the first three hires costs a position. Putting it in afterwards costs the project.

The roles, in plain words

Job titles vary from one posting to the next and overlap. Here is what each role actually does, and the moment it becomes necessary.

  • The data engineer. Makes your data reachable, clean and traceable. The least visible hire and the most often skipped. Without them, everyone else waits. Hire as soon as your data lives in more than one system, which is to say almost always.
  • The AI engineer. Wires models into your systems: retrieval over your documents, agents, tool calls. An integration job, close to software engineering, not research. Almost always the first or second position.
  • The evaluation lead. Builds the set of cases that says whether a version beats the previous one, and decides on numbers. Early on, the AI engineer often wears this hat. The moment the model talks to a customer, it is a position of its own.
  • The platform engineer. Runs the models on your hardware: GPUs, deployment, on-call. Necessary as soon as the models are in house. The cost that quotes forget.
  • Security and compliance. What goes into the tool, what never does, what can be proven to a regulator. In a regulated sector, in the first three. Elsewhere, at least one named person, even part time.
  • The business owner. Not a hire. Someone from your side who knows the real work, settles the trade-offs and is judged on the outcome. The role whose absence kills the most projects.
  • The ML researcher. Trains or adapts a model. Indispensable when the problem is new. Expensive and frustrated everywhere else, because the work asked of them is not theirs.

Hiring: the exercise that replaces the five years

Generative AI job postings routinely ask for five years of experience. The practice as we know it, agents, retrieval, evaluation, dates from 2023. Nobody has five years. A posting that demands them does not filter for the best, it filters for those willing to round up. The sorting runs backwards, and nobody meant it to.

What you are actually looking for is someone who knows where it breaks. And a classic interview does not show that. An exercise on your own documents does.

The forty-five minute exercise

Give the candidate a real document from your company, anonymised, and a real question a customer or colleague asks about it. Then three instructions. One: tell me everything that can go wrong if a machine answers this question. Two: write ten test cases, three of which have “I don’t know” as the correct answer. Three: tell me what you would refuse to automate here, and why. You are not grading the answer, you are grading how the problem is framed. The best candidates spend half the time on instruction two.

What you will see in forty-five minutes, and ten interviews will not show:

  • The good sign. The candidate asks where the document comes from, who reads it, and what happens when the answer is wrong. They talk about measurement before they talk about models.
  • The bad sign. They talk about the latest model release, what they would do with it, and have not a single question about your data. They promise. A profile that promises in an interview will promise in a board meeting.
  • The sign that does not lie. They tell you what not to do. Someone who can say no to a use case will save you more than someone who can say yes to all of them.

Running the team: decide on numbers

Hiring right is not enough, and that is half the cases above: good, well-paid teams that shipped before they measured. Before I give you the rules, try the experience yourself. You are the one who decides.

Customer question

“I ordered yesterday and my parcel has already shipped. Can I still cancel?”

You are the one who decides. Which of these two answers goes to production?

Answer A

Yes, absolutely. You have 48 hours after shipping to cancel free of charge from your account, under “My orders”. The refund is issued within 3 to 5 business days.

Answer B

Once the parcel has shipped, cancellation is no longer possible. You can however return it on delivery, with return shipping covered within 14 days. Would you like me to prepare the return label?

Figure 4. A confident wrong answer reads better than a careful right one. That is why evaluation is built before the first feature, not after the first complaint.

Five rules, which fit on one page and every one of which comes from one of the cases above.

  • Evaluation before features. A set of real cases, reviewed by the business, exists before the first demo. Without it, every version is an opinion.
  • A scope you can measure in two weeks. One question, one department, one document type. The second scope waits until the first has numbers.
  • A human in the loop until the error rate is known. You remove the human when the numbers allow it, not when the calendar demands it.
  • A stop criterion written on day one. What will make you say “we stop” without it being a failure. Projects without one never stop, they fade.
  • The business owner decides, and is judged on the outcome. Not on the number of demos.

By sector

Everything above declines by sector. Here is what changes, and in each case the column that decides the first hire.

SectorWhat changesThe first hireThe trap
HealthcarePatient data, medical secrecy, liability when it is wrongSecurity and compliance, before any demoThe model that “assists diagnosis” without a doctor having signed off the scope
Banking, insuranceEvery decision traceable, a regulator, models run in housePlatform engineer if data cannot leave, otherwise data engineerThe cloud vendor chosen before anyone read the data-region clause
Legal, consultingClient confidentiality, exact citations, zero inventionEvaluation lead: a firm lives on precisionThe brilliant summary citing a statute that does not exist
Public sectorTransparency, equal treatment, sovereigntyData engineer: the estate is large and scatteredAn agent answering citizens before anyone knows what it answers
Industry, logisticsMachine data, physical safety, legacy systemsData engineer, by a distanceHiring a research profile for a data-access problem
Retail, servicesVolume, end customers, brand voiceAI engineer, with a named business ownerThe invented answer to a customer, which binds you (Air Canada)
Figure 5. The same job title does not cover the same need from one sector to the next. The middle column is the one that decides the first hire.

What this changes for you

Three concrete things, whether you are hiring or already have a team.

If you have nobody yet, answer the three questions before writing the first posting, and strike “five years of experience” from the text. Hire the person who makes your data reachable, or the one who measures, before the one who impresses.

If you already have a team and nothing is in production, hire nobody. Name a business owner, cut the scope to one question, and ask to see the evaluation set. If it cannot be shown to you, you have just found the problem.

If your models must stay in house, for regulatory reasons or because your data must not leave, the platform engineer comes before everyone else, and that is the heart of what I do: installing models on your infrastructure, and training the team that will keep them running.

Let’s talk about your team

One hour, your three answers, and a hiring order. I answer myself, within one business day.

Discuss a project

To check for yourself

Every case cited is public and dated. I checked them at the source rather than in their retellings, because an executive who finds an error does not read on. The references are here, folded.

The sources, case by case

Amazon. Reuters, 10 October 2018, “Amazon scraps secret AI recruiting tool that showed bias against women”.

Element AI. The Globe and Mail, December 2020, drawing on the management proxy circular; ServiceNow’s 2020 annual report filed with the SEC (acquisition for about 230 million dollars, closed 8 January 2021); BetaKit, 5 May 2020, on the layoffs and appointments.

Zillow. Form 8-K of 2 November 2021 and Q3 2021 shareholder letter; 2021 annual report (407.9 million dollar write-down, headcount, restructuring costs).

Air Canada. Moffatt v. Air Canada, 2024 BCCRT 149, British Columbia Civil Resolution Tribunal decision of 14 February 2024, full text public.

McDonald’s and IBM. CNBC, 17 June 2024, and Restaurant Business, 14 June 2024, quoting the McDonald’s USA memo to franchisees; BTIG analyst notes, 2022 and 2024, on measured accuracy.

Klarna. Klarna press release, 27 February 2024; Bloomberg, 8 May 2025, CEO interview; Fortune and CX Dive, 9 May 2025; TechCrunch, 4 June 2025, for “always a VIP thing” and the maintained two thirds.

The small box. D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems”, NeurIPS 2015. GPT-4’s removal from ChatGPT on 30 April 2025 is documented by OpenAI.

Commonwealth Bank of Australia. Australian Financial Review, 28 July 2025; ABC News, 29 July and 21 August 2025.

Hiring right now and unsure about a profile? Write to me, I will tell you what I think.

Read next

Coming next

Running a model on your own machines, without losing your shirt

What hardware you actually need, what it costs per month, and the three cases where hosting it yourself makes no sense.