Anyone who says coding is solved has never inherited a legacy system. Or run production at scale. Or tried to explain to a product manager why their “small” AI-generated feature touches seventeen different services.
The “coding is solved” narrative has reached peak volume. CEOs are declaring pencils down on hand-written code. Influencers are telling engineers their “taste” is the only remaining value. And somewhere, a maintenance engineer is on-call at 3 AM debugging a stochastic model’s confidently wrong output.
Alex Ewerlöf, a veteran developer with two engineering degrees, posted a piece that racked up 417 points on Hacker News for calling this narrative exactly what it is: brain-dead. His argument deserves attention not because it’s contrarian, but because it surfaces the uncomfortable truth about what software engineering actually is versus what AI marketing wants it to be.
The Damning Evidence of “Solved” Coding
Let’s examine the track record of the companies most aggressively pushing the “coding is solved” line.
Anthropic owns the entire stack: the model, the harness (Claude Code), the prompts, and the runtime after acquiring Bun. According to Ewerlöf, they still shipped Claude Code with issues like:
– Returning Bun’s help menu instead of executing commands
– A CLI binary installer that silently deletes itself post-installation
– False rate limit errors despite available plan capacity
These aren’t edge cases from some random project. These are the foundational bugs in the very tool being marketed as a coding revolution. When the team with the most control and the deepest pockets ships this, the gap between the marketing and reality is staggering.
Even the status page shows the problem. As Ewerlöf points out with a screenshot, Anthropic’s incident history demonstrates that orange is the new green when it comes to their reliability track record.
Cognitive Debt: The Hidden Cost of Tab-Key Engineering
The Sogeti Labs research team has a term for what’s happening in organizations adopting AI coding tools: cognitive debt. Their analysis of the phenomenon is surgical.
GitHub and Microsoft’s early studies showed developers completing tasks up to 55% faster with AI copilots. Impressive. But the downstream effects tell a different story. Code exposure time in reviews is stretching. Architectural discussions are stalling. Reviewers are reporting mounting mental fatigue facing increasingly massive diffs.
This isn’t just anecdotal. GitClear’s analysis of over 150 million lines of code between 2020 and 2023 found that AI-generated code quality shows downward pressure, with code churn (code modified or reverted within two weeks) doubling compared to pre-AI baselines. The researchers described AI behavior like an itinerant short-term contractor: focused on completing the immediate task with zero regard for the long-term health of the codebase.
The mechanics behind this are subtle. An LLM generates code via contextual completion across a narrow perimeter, open files, immediate prompts, discovered imports. It optimizes locally at the expense of global coherence. Where an experienced engineer values restraint and reuses existing abstractions, the model takes the path of least textual resistance, producing self-contained, redundant code that passes tests but degrades the overall system.
The Reviewer’s Paradox
The most corrosive effect of AI-generated code is what it does to the review process. Traditional software development relied on cognitive symmetry: the mental effort of writing code organized the developer’s thoughts, so when a reviewer traced that logic, they followed something coherently designed.
Unchecked AI usage shatters this. Hitting Tab to accept a complex multi-line suggestion bypasses the deliberate design process entirely. The entire burden of critical inspection shifts to the peer reviewer, who now faces a diff that looks technically perfect, proper formatting, passing type checks, plausible docstrings, but may conceal domain logic flaws, resource leaks, or faulty concurrency assumptions.
Securing AI agents reveals the enduring importance of architectural boundaries. The same logic applies to reviewing AI output: the guardrails have to exist before the damage happens.
Cognitive Debt in Architecture Decisions
The problem isn’t limited to code level. A recent arXiv paper from the EuroPLoP 2026 focus group interviewed 22 practitioners from industry and academia about AI in software architecture. Their findings are sobering.
The researchers identified cognitive debt as a progressive erosion of human understanding when AI-assisted decisions are accepted without full intellectual engagement. This is the architectural version of the classic “it works, but I don’t know why” problem, except now it’s “it works, but nobody understands the system anymore.”
The focus group found broad consensus that architectural decision-making, accountability, and authoring architectural guardrails remain fundamentally human tasks. They introduced the concept of “harness engineering”, the discipline of building the systems that govern AI-assisted creation, including validation mechanisms, knowledge layers, and company-specific standards.
Why Functional Requirements Aren’t Solved Either
The “AI can write decent code” crowd misses something fundamental: creation was never the expensive part. As Ewerlöf puts it, anyone who has run software in production knows that maintenance, reliability, security, and scalability constitute the majority of cost. These non-functional requirements (NFRs) are where software either succeeds or collapses.
But even functional requirements, what the code is supposed to do, aren’t solved. The Dunning-Kruger effect is fully in play: people who don’t read AI output are the most confident in it. When I run agents to auto-generate massive PRs, the volume of code makes it humanly impossible to review carefully, so we either trust the output or dump it back.
The codes are deterministic. AI output is stochastic. Given the same input, a model can produce different outputs with the same probability distribution. We’re building platforms on foundations that can change their behavior without notification.
Real-World Complexity Is Not AI-Amenable
The mobility and fuel retail industry provides a perfect example. As MobilityPlaza’s analysis of legacy software complexity shows, the hidden operational costs in these systems are staggering.
Consider pricing: a single price change might affect the forecourt controller, price signs, card pricing, discount agreements, and invoicing. If these components don’t share the same data and rules, employees must manually verify the change propagated correctly.
Inventory management? Delivered volumes, tank measurements, and sales volumes come from different sources. Without automatic data reconciliation, wetstock management becomes a manual checking process.
Clearing? Cross-acceptance transactions from different networks, tariffs, commissions, and billing flows must be matched against supplier invoices.
Each of these seems manageable in isolation. Together, they create a landscape where technology complexity becomes operational complexity. The article makes a critical point: you pay for complexity twice, in operational hours and in commercial capacity. Every manual check and correction is time not spent on customers or growth.
The question every engineering leader should be asking is how many systems an employee needs to answer one customer question. That number is the true cost of your architecture.
The Economic Reality Check
There’s a deeper economic misalignment emerging. Software vendors have been charging human rates while paying for AI output prices. As Ewerlöf documents, this chasm is closing as more people wake up to the fact that they can create competent software at a fraction of the cost.
The two paths forward are stark: accept the price crash and lower quality accordingly, or maintain pricing by focusing on quality, using AI more thoughtfully, prioritizing understanding and accountability over velocity. The second path requires actual engineering skill, not just prompt engineering.
This isn’t about being anti-AI. It’s about recognizing that code is a side-effect of thinking and experimenting. Ewerlöf compares shrinking an engineer’s job to coding with shrinking a chef’s job to cutting vegetables. It’s part of the job, but it’s never been the end.
The Accountability Problem
If you ship a piece of code, you’re accountable for it regardless of how you produced it. This statement cuts through every excuse in the AI tooling playbook. AI cannot be held accountable. It can’t suffer consequences, pay fines, serve prison sentences, or care that it cost someone their job. The worst thing you can do to AI is unplug it.
The architecture focus group reached the same conclusion: accountability remains fundamentally human. The criticality criterion, the combination of uncertainty and cost of change, determines how much human oversight is needed. For critical systems, that oversight is absolute.
Frameworks for Managing the Mess
The Sogeti research proposes concrete safeguards that every engineering team should implement immediately:
Cognitive Complexity Gates: Enforce Campbell’s Cognitive Complexity metric as a blocking gate in CI/CD pipelines. Fail any PR introducing a method with a score above 10, forcing developers to refactor verbose AI outputs before requesting review.
Strict PR Size Limits: Cap PRs at under 200 lines of effective code (excluding test fixtures). This prevents dumping large, unbroken blocks of generated logic that nobody properly reviews.
Custom AST Rules Against Reinvention: Deploy internal semantic linters that flag generic utility definitions when standard company-wide libraries exist. This directly addresses the AI tendency to reinvent local date parsers, payload validators, and HTTP wrappers instead of using existing abstractions.
Valuing Code Deletion and Refactoring: Shift engineering culture to reward code deleted or simplified rather than raw lines committed. A senior developer’s primary value becomes acting as a critical curator.
What Actually Remains Human
The architecture focus group’s findings about agents solving already-solved problems while ignoring the actual mess ring true. They identified several areas where human involvement isn’t just preferred, it’s fundamental:
- Architectural decision-making
- Authoring architectural guardrails
- Determining criticality (uncertainty + cost of change)
- Validating AI suggestions against broader system context
The rate of broken ownership patterns is increasing. Removing people to rely on AI exposes hidden architectural dependencies in ways that surprise everyone. Using AI without understanding what you’re building isn’t leverage, it’s outsourcing the thinking that makes engineering valuable.
The “Taste” Problem
The most pernicious narrative of the AI era is that engineering has shifted to “taste”, that since AI can write the code, the only remaining human value is aesthetic judgment.
This is the lie retired chefs tell themselves.
Everyone has taste. It’s the rest of the discipline, knowing how things work, why they break, and how to fix them under pressure, that creates engineering value. AI has lowered the bar for generating decent-looking output while simultaneously raising the bar for what’s payable effort.
When a buyer purchases a SaaS product, they’re not paying for code generation. They’re paying for accountability, reliability, and a vendor who will answer when things break. As Ewerlöf notes, SaaS companies are increasingly in the business of selling SLAs. The competitive landscape is shifting to whoever can guarantee quality, not whoever can generate the most code.
Economic Realities of LLM Coding
The Sogeti research highlights another uncomfortable truth: when reviewing tools like Fable that operate without proper context and approval mechanisms, the output quality becomes questionable. Ewerlöf’s “Nordic Gold” analogy is apt: AI output looks like gold, costs a fraction of real gold, but doesn’t hold the same value under scrutiny.
The cost structure is inverted from what the narrative suggests. Code generation is cheap, but code understanding, the critical work of knowing what to build, why to build it, and what breaks when you build it wrong, is expensive. The total cost of ownership perspective shows that generation is only the tip of the iceberg.
The Fundamental Miss
The “coding is solved” crowd misses why languages exist in the first place. Graham’s Law, that natural language is vague and conflicting, is the primary reason programming languages were created. Compilers and type-checkers catch conflicts and syntax errors. When you ask an LLM to generate code from natural language, you’re introducing exactly the ambiguity that programming language designers spent decades eliminating.
The LLM can count the r’s in “Raspberry” wrong. It can suggest walking to the carwash when you asked about driving. These aren’t edge cases, they’re fundamental characteristics of stochastic text prediction. Wrapping them in harnesses and prompts and skills doesn’t solve the underlying logical unreliability, it just makes the output harder to trace.
For product management, the problem compounds. The “full spec upfront” idea fails because it’s impossible to meaningfully spec software ahead of time. Requirements emerge through iteration and experimentation. The engineering teams without product managers experiment showed this clearly: AI can’t replace product thinking because product thinking is an ongoing conversation with users, not a one-time specification.
What This Means for Engineers
If you’re an engineer reading this, you’re not obsolete, but your job description is changing. The engineers who will thrive are the ones who can work with AI while maintaining the understanding that makes their code trustworthy.
The recent benchmark results show AI models improving on specific, bounded tasks. But these tasks have clear specifications and test suites. Real software doesn’t work that way.
The future engineering roles emerging are:
– Technical product managers who translate business problems into buildable solutions
– AI deployment engineers focused on alignment, reliability, and governance
– AI quality engineers who tame stochastic outputs
– Harness engineers who build the systems that make AI output trustworthy
The Path Forward
This isn’t doom and gloom. AI is a bar-raiser. If your output quality equals or falls below AI-generated code, you need to upskill. But using AI doesn’t mean surrendering the fundamental disciplines that make software work.
The frameworks exist: cognitive complexity gates, PR size limits, architectural guardrails, and the discipline of understanding what you deploy. What’s missing is the cultural will to enforce them in the face of velocity pressure.
The next time a CEO declares that coding is solved, remember: the people making this claim have low risk tolerance for the software they depend on. Healthcare, finance, automotive, defense, aviation, these industries still need engineers who understand what they’re building. The plane that uses autopilot still has a pilot, and autopilot is a closed control system with deterministic behavior. AI isn’t there yet.
Keep your AI-free hobby project. Keep writing code. Keep understanding systems. In a world where output is increasingly free, understanding becomes the premium product. That’s the architecture no AI can generate.




