The bottleneck in AI hiring is no longer modelling ability. It is the far smaller number of people who have taken a model into production, watched it degrade, and had to defend its decisions to someone who could switch it off.
In fintech that last part is not hypothetical. If your model touches credit, fraud or any decision affecting a customer’s access to financial services, someone will eventually ask you to explain a specific outcome — and “the model said so” is not an answer.
First, work out which job you are hiring for
Most weak AI briefs are weak because they describe three roles at once.
Research / applied science. Novel modelling, experimentation, papers. Genuinely needed by a small number of companies, and by very few fintechs at Series A or B.
ML engineering. Trains, evaluates, deploys and monitors models. Owns the training pipeline, the feature store, drift detection and the retraining loop.
Applied AI engineering. Builds products on top of existing models. Owns retrieval quality, prompt and orchestration design, evaluation harnesses, guardrails, latency and cost.
Data engineering. Pipelines, warehousing, lineage and data quality. Frequently the actual constraint, and frequently the role that should have been hired first.
Writing “AI Engineer” across all four produces a shortlist where nobody is obviously right and the hiring manager cannot say why. Name the one you need.
What to screen for
Has anything they built reached production?
Ask for one specific model or system in live use. Then ask what degraded, and how they found out.
This is the most efficient filter available. Notebook-stage work and production experience look nearly identical on a CV and completely different in this conversation. Candidates with real production experience talk unprompted about monitoring, retraining and the gap between offline metrics and live behaviour.
Evaluation discipline
For any LLM-facing role:
“How did you know it was working? And how would you have known if a prompt change made it worse?”
You are looking for evaluation datasets, offline evaluation, human review loops and regression detection. Candidates who answer with “we tested it and it looked good” have shipped demos, not products. This has become the clearest dividing line in applied AI hiring.
Explainability, where it is required
In credit and fraud decisioning, a model you cannot explain is a model you cannot deploy. Ask directly:
“You have a model that outperforms the incumbent but you cannot explain individual decisions. It is going into a credit decision. What do you do?”
Candidates with regulated experience immediately discuss adverse action requirements, model governance, challenger models and the trade-off between performance and defensibility. Candidates without it often argue for the better AUC — which tells you exactly what you need to know.
Cost and latency awareness
Applied AI in production is an economics problem as much as a quality one. Ask what a system cost to run and what they did about it. Strong candidates talk about caching, model routing, prompt size and where they traded quality for cost deliberately.
What the pool actually looks like
Two things worth setting expectations on.
Deep production AI experience is younger than most job specs assume. Specs asking for eight years of LLM productionisation are describing a population that does not meaningfully exist. Look for strong engineering fundamentals plus demonstrable recent production work.
Fintech-specific AI experience is a genuinely small pool. If you require both deep production AI experience and prior regulated-decisioning exposure, you are searching a narrow market and should plan the compensation and the timeline accordingly — or decide which of the two you can train.
Sequencing mistakes
Hiring AI before the data platform. The most reliable cause of early attrition in this function. A strong AI engineer who cannot get clean, documented, timely data will spend a year on plumbing they did not sign up for, and then leave.
Hiring a Head of AI too early. At small scale, the need is usually someone senior who still builds. A leadership hire with no team, no platform and no mandate is an expensive way to discover that.
Screening LLM engineers on modelling theory. Asking someone who builds retrieval and evaluation systems to derive backpropagation tests a skill the job does not use, and rejects good candidates.
Using take-home exercises for senior AI candidates. This market is competitive enough that strong people decline them. Use a paid exercise or a live working session.
A process that works
- Screening call. Which of the four roles have they actually done? One production system, described specifically.
- Technical deep dive. Their system, their evaluation approach, what degraded and how they knew.
- Applied session. A live working session on a problem resembling your actual work — not a puzzle.
- Governance conversation. For regulated decisioning: explainability, monitoring, and what they would refuse to ship.
Building an AI or data team in fintech? Send us the brief, or read more about our AI and data recruitment.