September 19, 2026
23 min readThe Axis Moved: Specialist vs. Generalist in the Age of AI
Two interviews
I sat two technical interviews in the space of a single year, and the distance between them is the reason I am writing this.
The first was with an enterprise software company that was still early in its own adoption of AI. The questions were detailed and specific. What exactly had I used, how exactly had I configured it, what did the implementation look like at the level of the individual setting. It was a good interview, run by good engineers. It was also, in hindsight, a recall test with a long tail. The more precisely I could describe the tools I had touched, the better I did.
The second was with an AI native company. On paper the questions looked almost generic, the sort of thing you might mistake for warm up conversation if you were not paying attention. Why would you choose this over that. What breaks first when the traffic goes up. Which algorithm, and what does it cost you. How would you draw this system if you had to hand it to someone else. There was nothing to recall. Every question was asking the same thing underneath: do you actually understand why any of this works.
Same candidate, same year. Two very different definitions of qualified.
I walked out of the second one convinced I had just seen where technical hiring was going, and completely wrong about what it meant. Sorting out how I was wrong took me through a fair amount of labour market data, and the rest of this post is what I found: the axis I had been reading the market on for years, the reason AI looked like it flipped that axis, the evidence that it did not, and what I think actually moved.
The axis I used to believe in
For most of my career I read the job market on a single axis, and I was confident about it.
In Türkiye, I thought, you win by being a generalist. Teams are smaller, budgets are tighter, and the person who can move between the backend, the database, the deployment pipeline and the occasional production fire is the person who keeps the product alive. Nobody is paying for depth in one narrow layer when there is nobody else to cover the other layers.
In Europe and the United States, I thought, you win by being a specialist. That was not a feeling. It was what the listings said. When I applied to roles abroad, the descriptions named the layer, the stack, and often the exact library, and the screening questions followed the description closely. When I applied at home, the same seniority level came with a much wider surface and a much looser description.
I want to be careful here, because this is the least defensible paragraph in the post. This is one engineer’s sample, gathered over a limited window, in a limited set of domains, filtered through the roles I happened to find interesting. It is not a claim about two labour markets. It is a claim about what I saw. I saw it consistently enough to build a career strategy on it, which is a different thing from it being true in general.
Then the tools arrived, and the axis I had built everything on stopped making sense.
Why AI looked like it flipped that axis
The obvious thing happened first. Writing code got cheap.
Not free, and not always good, but cheap enough that the ability to produce working code stopped being a useful way to tell people apart. A designer ships a web app. A backend engineer ships a mobile client. I open a language I have never written professionally and I am productive the same afternoon. Whatever moat existed around knowing a particular framework well drained out in about eighteen months.
If that were the whole story, my conclusion would write itself. If every specialist can now do every other specialist’s job, specialisation is worth nothing, and the generalist wins by default.
It is not the whole story, and the gap shows up fastest in security.
There is a moment in a Turkish YouTube walkthrough of vibe coding that I keep returning to, because it names the problem in a single sentence. The presenter stops and asks: “Biz işte cloud koda uygulamamı güvenli yap dediğimizde neden güvenliği yapamıyor acaba?” Roughly: why is it that when we tell the model to make our application secure, it cannot actually do the security [7].
That question is this entire post in miniature. He is not asking whether the model can write code. He knows it can, he just watched it. He is asking why the instruction that any reasonable non expert would assume is sufficient turns out not to be.
The measurements answer him. Veracode has now evaluated more than 150 models on tasks that can be completed either securely or insecurely. Syntax correctness runs above 95%. The security pass rate sits at roughly 55%, which is almost exactly where it sat two years ago. Close to half of generated code introduces a known vulnerability when no explicit security guidance is given, and making the models bigger has not moved that number [4].
Apiiro’s analysis of repositories inside Fortune 50 companies shows what this looks like once it reaches real production code. Developers using AI assistants produced three to four times more commits. Syntax errors dropped 76% and logic bugs dropped 60%. Over the same period, privilege escalation paths rose 322% and architectural design flaws rose 153%. Their one line summary is better than anything I could write: AI is fixing the typos but creating the timebombs [6].
Read those two findings together and the shape is hard to miss. The machine became excellent at the layer where correctness is local and checkable. It did not improve at the layer where correctness depends on knowing what the system is for, who is permitted to do what, and what happens at the boundaries. Those are not coding problems. They are design problems wearing a coding problem’s clothes.
Hiring noticed. Reporting on 2026 interview trends describes a Google pilot that adds a code comprehension round in which candidates work through an existing codebase with a company provided Gemini assistant, with interviewers scoring AI fluency directly: prompt quality, output validation, debugging. The framing is human led, AI assisted. Meta is described as running an AI enabled coding round where candidates can switch between models and are scored on problem solving, code quality, verification and communication. I have not been able to tie either of those to a primary company source, so take them as the trade press reports them rather than as documented policy.
It is worth being accurate about the scale of this, because it is easy to overstate. These are pilots and specific rounds, not company wide policy. The same reporting suggests most companies still prohibit AI during interviews, with roughly 38% permitting it in the United States, and some large employers, Amazon among them, holding a firm no AI line. The direction is clear. The arrival is not.
Still, look at what the AI enabled rounds actually measure. Not recall. Verification, judgement, and the ability to explain why the generated answer is wrong. That is a deliberate move away from testing whether you can produce code and toward testing whether you can evaluate it.
So here is the conclusion nearly everyone reaches, and I reached it too: if syntax is commoditised and judgement is not, then breadth beats depth, and the generalist finally wins.
It is a good argument. The data does not support it.
Except the data disagrees
I have to stop and be honest here, because when I started writing this post my conclusion was already written. Generalists win. I believed it before I checked it, which is exactly the wrong order.
Then I went looking for evidence, and the best evidence available says the opposite.
Revelio Labs analysed 75 million tech job postings collected since 2024. The average number of skills listed per posting fell by about 25%, from 30 down to 21. Over the same period, the average required years of experience attached to those skills rose by roughly 4%. Their own reading is direct: this indicates more demand for depth and expertise, and less demand for generalists. What is being asked for moved the same way, away from broad categories like software development and toward narrow ones like Java API development [1].
PwC’s 2026 Global AI Jobs Barometer, built on more than a billion job advertisements across 27 countries, points the same direction from the money side. Hiring for AI specialist roles grew 68.9% year over year. Jobs requiring AI skills are growing roughly eight times faster than the overall jobs market, 69% against 9%, and the wage premium for AI skills reached 62%, up from 57% the year before [2].
So if the question is whether companies are asking for more specialists or more generalists, the answer written on the page is: specialists. Not ambiguously, and not narrowly.
I could have left this section out. The post would have been shorter, cleaner, and considerably more fun to write. But then I would have been describing the market I wanted rather than the one that exists, and anyone can check Revelio’s numbers in about ninety seconds.
The naive version of my thesis is dead. What interests me is what survives it.
What actually moved
Go back to the Revelio finding, and this time read the second half of it.
Postings are asking for fewer skills and more years of experience. Under the old frame those two should not move in opposite directions. A shorter skill list is what specialisation looks like, fine. But a specialist is not automatically a senior. That rising experience requirement is measuring something separate.
Revelio isolates it. Controlling for seniority, industry, role and description length, postings that call for AI usage list about 3% fewer distinct skills and require about 11 more days of experience. Their explanation is the part worth sitting with: companies are outsourcing skill sets they used to demand from workers, and in exchange they want the expertise and judgement that come with more experience, plus whatever is not comfortably delegable to AI [1].
Read that as a sentence about pricing. The skills line got shorter because skills got cheaper. The experience line got longer because judgement did not.
PwC finds the same thing from the opposite end of the career ladder. In US data, entry level roles in the most AI exposed occupations are seven times more likely to demand traditionally senior level skills, specifically judgement and leadership. Those seniorised entry level roles grew 35% since 2019, while other entry level roles fell 10% [2].
And in January 2026 the IMF put the overall shape of it into a staff discussion note. Across millions of postings, demand for new IT and AI skills raises average wages and employment, but deepens polarisation, with the gains concentrating among high skilled workers and, through higher consumption of services, low skilled workers, potentially contributing to a shrinking middle class [3].
Three datasets, three methods, one shape. The market is not choosing breadth over depth, or depth over breadth. It is hollowing out the middle and repricing judgement at the top.
Which is why the generalist and specialist question stopped predicting anything. It was not the wrong question because the answer flipped. It was the wrong question because it measures the wrong dimension.
Here is how I would put the new question, and I think it is a better question to be asked:
It used to be “how many technologies do you know?” It is becoming “how much of a decision can we hand you, and trust you to see it through from the first why to the last consequence?”
I find that genuinely encouraging, because it means the thing being rewarded is something you can grow rather than something you have to keep re-learning. It also rescues almost everything I believed and kills the part that was wrong.
What I had been calling a generalist was never really breadth of technology. When I claimed generalists adapt faster, I was not claiming they had memorised more frameworks. I was claiming they understood why systems are built the way they are, so that a new stack became a translation problem rather than a learning problem. That is not breadth of tools. It is breadth of judgement. I had the right instinct attached to the wrong noun, and the wrong noun is what made the claim falsifiable.
It also explains which specialists are genuinely struggling, because it is not all of them.
Depth in an implementation surface is what AI absorbed: knowing a framework’s API exhaustively, knowing the idiomatic pattern, knowing the quirks of a particular ORM. Depth in a problem domain is what it did not: knowing why this system needs idempotency, what breaks at ten times the traffic, which consistency guarantee the business can actually afford to lose.
Both of those used to look identical on a CV. They do not pay the same any more.
Apiiro reports a case that makes the distinction concrete. An AI driven pull request changed an authorization header across several services. A downstream service was not updated to match, and the result was a silent authentication failure [6]. Every individual file in that change was almost certainly correct. The diff would have compiled, passed review at a glance, and satisfied any linter you pointed at it. The failure lived in the space between the services, which is the one place no file level check was looking.
Now imagine two engineers reviewing that pull request. The first knows the framework deeply and reads each file carefully. Every file is fine, so the review passes. The second does not necessarily know the framework better, but knows what the system is for, and asks who else reads this header. That question takes four seconds and is worth more than the entire rest of the review.
That gap is the whole argument. It is not breadth versus depth. It is whether your knowledge is shaped like a file or shaped like a system.
So the losing profiles sit on both sides of the old axis. The shallow generalist who knows six frameworks badly is competing directly with a tool that knows sixty. The deep specialist whose depth lives in an implementation surface is in the same position, just with a better looking résumé. The profile that is gaining is the one that can hold the whole decision, and that was never a question of how many languages you list.
The map
If I try to be systematic about this, my instinct is to walk the software lifecycle and ask at each stage what AI actually took and what it left behind. Here is my working map. I consider it a v1 and I fully expect to revise it.
| Stage | What it used to require | What AI absorbed | What it did not, and why | Grounded in |
|---|---|---|---|---|
| Requirements and scoping | Turning a vague ask into something buildable | Drafting, structuring and rewriting the spec once intent is clear | Deciding what the system is for, and what you are willing not to build | PwC on entry level roles now demanding judgement and leadership [2] |
| Architecture and system design | Knowing the patterns and being able to justify a choice | Producing a plausible design for a familiar problem | Choosing which failure mode you can afford, which is a business question before it is a technical one | Apiiro: architectural design flaws up 153% [6] |
| Implementation | Fluency in a language, its framework and its idioms | Most of it, at very high syntactic accuracy | Honestly, very little. This is the layer that moved | Veracode: syntax correctness above 95% [4]; Apiiro: syntax errors down 76%, logic bugs down 60% [6] |
| Code quality and conventions | Internalised house style and review taste | Local consistency and mechanical conformance | Knowing which conventions encode a real constraint and which are just habit | Guides and sensors framing from my harness engineering post [8] |
| Security and threat modelling | Knowing the common vulnerability classes and where they hide | Very little that is reliable without explicit guidance | Modelling who the attacker is and what they want, which is context the model does not have | Veracode: security pass rate around 55%, flat for two years [4]; Apiiro: privilege escalation paths up 322% [6] |
| Verification and review | Writing tests and reading diffs | Generating tests and summarising diffs | Deciding what done means, and noticing what is absent rather than what is wrong | Apiiro: larger pull requests overwhelming review [6] |
| Operations and incident response | Runbooks, dashboards, pattern recognition under pressure | Triage assistance and hypothesis generation | Holding the blast radius in your head while the system is on fire | My own experience, and the harness argument for state and verification [8] |
The column that matters is the fourth one. Read it top to bottom and none of the entries are about a language, a framework or a tool. They are all about knowing what the system is supposed to do.
What I would actually do about it
If you are an engineer making a career bet.
Stop counting technologies and start counting decisions you have owned end to end. A CV listing nine frameworks now reads as nine things a model also knows. A CV describing three systems you designed, with the trade offs you made and the one you got wrong, reads as something else entirely.
Pick a domain rather than a stack. Payments, fraud, distributed storage, aviation systems, whatever it is. Depth in a problem domain compounds. Depth in an implementation surface depreciates, and it depreciates on the release schedule of somebody else’s model.
Get deliberately good at verification. The scarce skill in the AI enabled interview rounds is the same scarce skill on the job: looking at plausible generated output and knowing what is missing. That is a skill you can practise directly, and almost nobody does.
I am not alone in thinking this. Thoughtworks ran two retreats on the future of software engineering this year, one in Utah in February hosted by Martin Fowler and one in Engelberg in June, and the findings from the second one land on exactly this point. Across sessions on testing, legacy modernisation, code review and team design, the same conclusion kept surfacing: code generation is no longer the bottleneck, verification is, because agents can produce code, specs, tests and infrastructure far faster than any team can trust it. Their recommendation is a renewed focus on cheap, fast, human legible verification, including characterisation tests, mutation testing and production back testing, and a hard look at the assumption that a manual code review is a guarantee of quality [9]. If you want to know where to spend your next six months, that list is not a bad place to start.
Make your judgement visible. Judgement leaves no artifact by default. Write architecture decision records. Write up the incident. Write the post explaining why you chose the boring option. If you cannot show the reasoning, the market has no way to price it, and it will price you on your skill list instead.
And do not skip the fundamentals on the grounds that the model covers them. You cannot verify what you do not understand. That is the honest answer to the question in that video: the reason “make it secure” does not work is that the instruction contains no information. Security is a set of decisions about a specific system and a specific adversary, and if you cannot make those decisions, you also cannot tell whether the model made them.
If you are a manager writing job descriptions and interview loops.
Your job description is a filter you probably wrote by accident. If it lists twenty one skills, you are screening for recall, and you will reliably select the candidate who is best at matching keywords rather than the one who is best at deciding things.
Interview for verification explicitly. Hand candidates generated code with a real, non obvious flaw in it, ideally one that lives between components rather than inside a function, and watch what they do. It is a better signal than any algorithm question I have ever been asked or asked anyone else.
Stop treating years in framework X as a proxy for judgement. It was always a weak proxy. It is now close to meaningless, because the framework specific part is exactly what got automated.
And there is one problem I do not have a clean answer to, so I will state it rather than dress it up. The routine work is how people used to become senior. If you remove all of it, you remove the training path. PwC’s entry level findings are a warning about this and not a celebration of it.
The Thoughtworks retreats spent real time on this, and their conclusions are more nuanced than mine, so I would rather pass them on than flatten them. The Utah retreat pushed back on the idea that AI eliminates the need for juniors at all. Juniors, they argued, are more profitable than they have ever been, because AI gets them past the initial net negative phase faster and they adopt the tools more readily than seniors do. The population they worried about was mid level engineers who came up during the hiring boom and may not have the fundamentals to thrive in the new environment, which is most of the industry by volume, and which nobody has worked out how to retrain [10]. The Engelberg retreat then raised the other side of it: if senior engineers pair only with agents, juniors lose the hands on path to judgement, taste and production instinct that the industry has always relied on to grow its next seniors. Most participants agreed that people building software this way need to know what good looks like, which is why seniors are thriving, and that this is precisely what makes the apprenticeship problem harder rather than easier [9].
Notice how neatly that matches the polarisation story from the labour market data. The people at the top thrive because they can tell good from plausible. The people at the bottom adapt because they have nothing to unlearn. The people in the middle are the ones the axis moved underneath. Whoever works out how to grow senior engineers in an environment with less routine work will have solved something genuinely important. As far as I can tell, and as far as the people in those rooms could tell, nobody has.
Honest caveats
A few things I want to put next to the argument rather than bury underneath it.
This is a claim about one era, and there is a named mechanism for it inverting. Ide and Talamàs model AI inside a knowledge hierarchy and find that non autonomous AI, the copilot, disproportionately benefits the least knowledgeable, while autonomous AI primarily benefits the most knowledgeable. Their model goes further: basic autonomous AI pushes humans toward complex problem solving, while advanced autonomous AI reallocates humans back toward routine knowledge work [5]. If that second transition arrives, much of what I have written here inverts, and the thing I am telling you to invest in becomes the thing that gets reorganised. I do not know when or whether that happens. I only know it is not idle speculation.
Two of my best sources sell security products. Veracode and Apiiro both benefit commercially from the finding that AI generated code is insecure. I used them because the methodologies are public and the numbers are consistent with each other and with independent academic work, but you should discount them accordingly, and I would rather say that than have you notice it yourself.
The interview round details are secondhand. The Google and Meta examples, and the 38% figure, come from trade reporting I have not been able to trace to a primary company source. I have left them in because the direction they describe is corroborated by everything else here, but they carry less weight than the numbered sources below.
Job postings are not hires. Everything in the data sections describes what companies advertise. What they advertise and what they sign are different documents, and the gap between them is not measured here.
My Türkiye and Europe observation remains anecdotal. I said it once already and I will say it again at the end, because it is load bearing for the narrative and unverified as evidence.
Aggregate data does not describe your career. Every number in this post is a distribution. You are one observation. A well positioned specialist in a domain with structural demand is doing extremely well right now, and no labour market average tells you otherwise.
Closing
I set out to argue that generalists win. I think the instinct underneath that was right and the sentence was wrong, and the distance between the two turned out to be the interesting part of the problem.
The market is not choosing breadth over depth. It is paying for the part of the work that does not compress: knowing what the system is for, what it must never do, and what you are prepared to trade away to get it. Everything around that is getting cheaper every quarter. That part is not.
Which brings me back to the two interviews. The first asked what I had used. The second asked what I would decide.
At the time I thought the second company had simply modernised its process. I think now they were pricing the thing that was about to become scarce.
References
- Dean Boerner, Tech Hiring Is Getting More Specialized: What the 2026 Tech Job Market Wants, Revelio Labs, Sep 1, 2026. https://www.reveliolabs.com/news/tech/tech-hiring-specialized-job-market-2026
- PwC, 2026 Global AI Jobs Barometer, Jun 15, 2026. https://www.pwc.com/gx/en/news-room/press-releases/2026/pwc-2026-ai-jobs-barometer.html
- Jaumotte, Kim, Koll, Li, Li, Melina, Song and Mendes Tavares, Bridging Skill Gaps for the Future: New Jobs Creation in the AI Age, IMF Staff Discussion Note SDN/2026/001, Jan 2026. https://www.imf.org/en/publications/staff-discussion-notes/issues/2026/01/09/bridging-skill-gaps-for-the-future-new-jobs-creation-in-the-ai-age-572136
- Veracode, Spring 2026 GenAI Code Security Update and 2026 GenAI Code Security Report. https://www.veracode.com/blog/spring-2026-genai-code-security/
- Enrique Ide and Eduard Talamàs, Artificial Intelligence in the Knowledge Economy. https://arxiv.org/abs/2312.05481
- Itay Nussbaum, 4x Velocity, 10x Vulnerabilities: AI Coding Assistants Are Shipping More Risk, Apiiro, Sep 2025. https://apiiro.com/blog/faster-code-greater-risks-the-security-trade-off-of-ai-driven-development
- Mesut Çevik, vibe coding walkthrough, YouTube, at 11:03. https://youtu.be/_X7vw-1XKrU
- Mustafa Arslan, Harness Engineering: How We Got From Prompts to Here, Aug 12, 2026. https://mustafarslan.me/blog/harness-engineering
- Thoughtworks, The Future of Software Engineering: Retreat Findings, Engelberg, Switzerland, Jun 28 to 30, 2026. https://www.thoughtworks.com/content/dam/thoughtworks/documents/report/tw_future_of_software_engineering_europe_2026.pdf
- Thoughtworks, The Future of Software Development Retreat: Key Takeaways, Utah, Feb 2026. https://www.thoughtworks.com/content/dam/thoughtworks/documents/report/tw_future _of_software_development_retreat_ key_takeaways.pdf
Click any figure or table to zoom.