Around 2011, José Valim, creator of Elixir, made a decision that had nothing to do with artificial intelligence. He chose explicit over implicit, and then he kept choosing it. Not as a slogan on a wiki page. As a rule applied over and over, in places where the implicit version would have been shorter and cleverer and easier to demo.
I’ve long suspected that José is a machine.
Fifteen years later, agents write Elixir with unusual accuracy. A Tencent study found that 97.5 percent of Elixir problems were solved by at least one model, the highest of any language tested, with Claude Opus 4 scoring 80.3 percent against 74.9 for C# and 72.5 for Kotlin. José wrote up the reasons in February, and they are almost entirely decisions made for human benefit that happened to pay off for machines.
Nobody planned that. The other machines just needed the same thing we did.
That explanation is too easy
If explicitness were the whole story, this would be a short article. Be more explicit. Write more down. Problem solved.
It doesn’t work, and the reason it doesn’t is that implicit doesn’t announce itself. Nobody sits down and decides to leave an assumption unstated. It accumulates. It hides in the gap between what your code says and what your team knows, and it wears a different disguise depending on where it’s hiding.
So the useful question isn’t whether to be explicit. It’s where implicitness gets in, and that turns out to have exactly three answers, because an agent has exactly three inputs.
There’s what it knew before it ever saw your code. There’s what your code shows it. And there’s whatever tells it that it got something wrong. Priors, evidence, correction. Nothing else is in the loop.
The wheels come off when all three fail at once. That’s why some stacks tolerate an enormous amount of unguided generation and others fall apart in a month. It isn’t one property. It’s a product.
Priors: what it brought with it
The model arrived already knowing how to write your language, which is the problem.
It learned from a corpus, and corpus quality varies enormously. Some languages are represented by fifteen years of production code written by people who cared. Others are represented by the most-copied Stack Overflow answer of 2016, reproduced ten thousand times by people in a hurry. The agent has no way to weight these and no reason to.
It also learned from every version of every library that ever shipped, blended together with no timestamps. So it writes Phoenix 1.6 idioms into a 1.8 application, confidently, because both were true once and nothing in your file says which year it’s living in.
And it optimizes for the common case rather than the correct one. Ask for a solution and you get the median answer from the corpus, which is not the best answer and is frequently not even a good one. In a language with thin coverage you get worse: it invents, fluently.
Its priors are the internet. You can’t change them. You can only override them, and you have to know they’re there to try.
Unless you’re an ecosystem, and you’re patient. Elixir keeps documentation separate from code comments and runs the examples in that documentation as tests, so a doc that lies fails the build. The language has been on 1.x since 2014, which means there’s far less contradictory history for a model to blend together. Neither of those choices was made for the benefit of a machine. Both of them produced a cleaner corpus anyway, over a decade, in public.
Evidence: what your code shows it
This is the bucket you actually control, and the one most teams have never thought about.
Mutability hides where state changes, so a function that looks pure is being edited from three rooms away. Runtime topology barely appears in source at all: supervision trees, process boundaries, what’s linked to what, which failures take down which neighbors. The agent reads modules. Your system is processes. It will write something correct as a function and wrong as a process, and nothing in the file it’s reading would have told it otherwise.
Then there’s everything your team knows and never wrote down. Soft deprecations, where a thing still works and everyone senior knows not to use it. Optional interfaces, where two paths are both supported and one is clearly preferred by people who’ve been here a while. Dead code, sitting there looking like an endorsement. Conventions that live in the reviewer’s head and surface only as comments on a pull request.
The obvious reading is that some languages hide more than others. That’s not quite it. What matters is whether there’s a marked path from the surface to what’s underneath. use Foo in Elixir means Foo.__using__/1. Always. It’s greppable, documented, and the agent knows the convention. That’s metaprogramming you can follow. Compare method_missing, or a runtime proxy assembled from annotations at startup. Same nominal problem, no path, good luck.
Correction: how it finds out it’s wrong
The third input is everything that argues back.
Compiler errors. Exhaustiveness warnings when you didn’t handle every case. Static analysis. Failures that happen loudly and near their cause instead of quietly and three layers away. Every one of these is a channel through which the system can tell an agent it’s wrong without a human being involved, and the number of those channels varies wildly by stack.
Two of them deserve more suspicion than they get.
Tests written by the same agent that wrote the code certify nothing. They encode the implementation’s assumptions, then confirm them. Green means the thing agreed with itself.
And a team where human review is the only real correction channel doesn’t have a system, it has a bottleneck. That was survivable when code was expensive to produce. It isn’t now.
🎯 Join Groxio's Newsletter
Weekly lessons on Elixir, system design, and AI-assisted development — plus stories from our training and mentoring sessions.
We respect your privacy. No spam, unsubscribe anytime.
Building the foresight in
Here’s the move that changes the arithmetic, and it’s the most useful thing in this article.
Good tooling converts priors problems into evidence problems. Version skew looks unfixable, since you can’t retrain the model. But mix.lock states exactly what you’re running, and version-specific documentation is available to the agent while it works. The prior is still wrong. The ecosystem just handed it a way to check.
That’s worth generalizing past Elixir, because it’s a design goal you can adopt for your own applications. Assume something will read this code that wasn’t in the room when you built it, and that has no access to anyone who was. What would it need in order to verify its own assumptions? Which decisions would it have to guess at, and where would you rather it didn’t?
Most teams have never designed for that reader. It’s arriving anyway.
What all of it has in common
Implicit is where the wheels fall off. If it announced itself this would be an easy problem, and it isn’t, because implicit wears a different disguise in every bucket. A version nobody stated. A convention nobody wrote down. A preference nobody marked. A failure nobody hears.
José is right again. He picked explicit over implicit for human readers, fifteen years before any of this existed, and it turns out to be the foundation of good agentic coding.
Trust but verify is cheap when your assumptions are on the surface. Everywhere else, you’re just trusting.
Teaching the language skills and prompting techniques that keep the wheels on is cheaper than a rescue. It starts with a call, not a pitch. Reach out to us at grox.io.
🛠 Keep the Wheels on with Structured Team Training
This post is from Bruce Tate's series on what the AI coding crisis is doing to engineering teams — and what it would take to train through it instead of around it. Groxio runs private training and ongoing advisory for engineering teams using AI with Elixir, Phoenix, OTP, LiveView, Ecto, Ash, and Postgres. We start with a diagnostic conversation about where your review queue, your seniors, and your codebase actually are.
— Bruce