Everybody Hired an AI Engineer. Nobody Knows What They Actually Do.

The job title did not exist at scale three years ago.
Then it did. Everywhere. Simultaneously.
AI Engineer. Machine Learning Engineer. Generative AI Specialist. Applied AI Lead. The names varied. The salary bands were remarkably consistent. Three hundred thousand dollars. Three fifty. Four hundred at the top end in San Francisco.
Every company needed one. Every company was competing for the same small pool of people who had shipped something real with LLMs. The bidding was frantic.
Fast forward to today.
The companies that won those bidding wars are sitting on eighteen months of results. The honest ones are asking a question that is uncomfortable to ask out loud.
What exactly did we hire?
The role that meant five different things
Ask ten companies what an AI engineer does and you get ten different answers.
At some companies, the AI engineer fine-tunes models. They manage training runs, curate datasets, evaluate model performance, and own the model artifacts that the company’s products are built on.
At other companies, the AI engineer builds RAG pipelines. They index documents, manage vector databases, tune retrieval, and make sure the model has the right context for the right query.
At others, the AI engineer writes prompts. They run experiments to find the system prompt configuration that produces the best outputs. They own the prompt library. They iterate on instruction tuning for the company’s specific use cases.
At others still, the AI engineer is a product engineer who happens to work on AI features. They call APIs. They build the application layer. The AI is the domain, not the discipline.
All of these are real jobs. None of them require the same skills. None of them produce the same output.
The companies that hired AI engineers mostly did not know which of these they were hiring for.
They knew they needed AI. They hired someone called an AI engineer. They figured the details would work out.
The details did not always work out.
The supply that materialised from nowhere
Here is what made the hiring wave so strange.
In early 2023, genuine AI engineering experience was extremely scarce. The people who had actually shipped LLM-powered products in production could be counted in the thousands globally.
By late 2024, there were hundreds of thousands of people describing themselves as AI engineers on LinkedIn.
The supply did not come from universities producing new graduates with AI engineering degrees. It came from the existing developer population rebranding.
Backend engineers who had spent three months working with the OpenAI API became AI engineers. Data scientists who had integrated LLMs into their workflows became AI engineers. Frontend engineers who had added a chat interface to a product became AI engineers.
Some of these people were genuinely excellent. The skills transferred. The rebranding reflected real capability growth.
Some of them had learned enough to pass the interview. The rebranding reflected the salary differential more than the capability.
The companies hiring at speed, under competitive pressure, paying four hundred thousand dollars for roles they barely understood, were not well-positioned to tell the difference.
The interview that did not catch it
The AI engineering interview in 2024 converged on a specific format extremely fast.
Build a RAG pipeline. Implement a simple agent. Fine-tune a small model on a toy dataset. Discuss evaluation metrics. Walk through a system design for an AI-powered feature.
These assessments tested whether candidates could do the specific things the interviewer thought AI engineering involved.
They did not test whether candidates had the judgment to know when to use AI and when not to. Whether they could debug a production AI system that was behaving unexpectedly. Whether they understood the operational reality of running inference at scale. Whether they could push back when a product requirement needed AI in a way that would not work.
The interviews selected for candidates who had practiced the interview format. The format was well-known because it had spread rapidly through the community.
The skills that determine whether someone is actually good at building AI systems in production are harder to test in an interview and were largely untested.
What eighteen months revealed
The companies that hired well know it now because they can see what their AI engineers built.
The products are in production. The systems are running. The AI engineer’s fingerprints are on architecture decisions that are either paying off or creating problems. The calibration between what was promised in the interview and what was delivered in the work is visible.
The companies that hired less well also know it now.
The AI features shipped but retention is soft. The pipeline architecture that seemed reasonable at design time has accumulated technical debt that is expensive to pay down. The production AI system generates costs that were not anticipated because the engineer who built it did not have deep enough experience with inference economics.
The eighteen months have been the world’s largest A/B test on what AI engineering talent actually is and whether the interview process selected for it.
The results are not uniformly flattering.
The companies that got it right did something different
The organisations with the best outcomes from their AI engineering hires shared a practice that most companies skipped.
They knew precisely what they needed before they hired.
Not “we need AI capability.” Something specific. We need someone who has run model evaluation pipelines at production scale. We need someone who has debugged latency issues in multi-step agent systems. We need someone who has made the build-versus-buy decision on model infrastructure and lived with the consequences.
Specificity in the requirement produced specificity in the hire. The interviews were designed around the specific problems the company actually needed to solve. The candidates who looked good were the ones who had solved those specific problems before.
The companies without that specificity hired for the general shape of the role and got a range of outcomes that reflected the range of what “AI engineer” actually meant.
The salary that is already correcting
The four hundred thousand dollar AI engineer salary was a product of a specific moment.
Scarce supply. Voracious demand. Every company bidding against every other company for the same small pool of people who had shipped real AI systems.
That moment has passed.
The supply has expanded. The companies that hired aggressively are now digesting those hires rather than adding more. The economic environment has made CFOs more attentive to whether expensive headcount is producing proportionate value.
The correction is not dramatic. The best AI engineers, the ones who have built real systems, debugged real failures, and made real architectural decisions that held up in production, are still commanding strong compensation.
The correction is hitting the middle of the distribution. The engineers who were priced like the best because the market could not distinguish them from the best.
The market has had eighteen months of signal now. It is updating.
What the role actually is
The most useful framing for AI engineering that has emerged from the hiring wave is not a job title.
It is a skill combination that does not map cleanly onto existing disciplines.
The AI engineer who produces the most value is a software engineer who understands production systems deeply, can reason about model behavior empirically rather than theoretically, has strong intuitions about when AI adds value and when it adds complexity, and can debug the specific failure modes that appear when you put probabilistic systems into production environments designed for deterministic ones.
This is not someone who has memorised the transformer architecture. It is not someone who can implement backpropagation from scratch. It is someone who has shipped things, watched them break in unexpected ways, and developed the judgment that comes from that cycle repeated enough times.
That person was rare in 2023. They are less rare now. Eighteen months of companies running real AI systems in production has produced genuine experience at scale.
The next wave of hiring will be better than the last. Not because the interviews will improve.
Because the people being interviewed will have actually done it.
The title created the demand. The demand created the experience. The experience is now real.
And the companies that understand what they are looking for this time will find something that the companies last time mostly did not.
The actual thing.