I’ve built developer tools for most of my career—CLIs, internal SDKs, deployment scripts, local simulators—and I keep running into the same pattern. The teams building these things get treated like a service desk, not an engineering discipline. The tools themselves get cobbled together in a rush, handed off with zero documentation, and abandoned the moment the original author moves on. Nobody budgets time for maintenance, nobody writes tests, and nobody thinks about the developer experience as something that actually matters. We end up shipping half-finished tools to our own colleagues and then act surprised when velocity stalls.
This isn’t a small problem. It’s a cultural failure that costs real money, burns out good engineers, and slows down entire organizations. The way we treat internal dev tools reflects a deeper assumption: that tooling work is somehow less important than product features. It’s time we call that assumption what it is—wrong.

Why We Keep Getting This Wrong
The root cause is simple but stubborn: most engineering organizations measure success by customer-facing output. Features, uptime, revenue impact—those get tracked, celebrated, and rewarded. Internal tools, by contrast, have no clear owner and no clear metrics. They’re the “extra” work that somebody picks up between sprint commitments, usually with a shrug and a note to “just get something working.”
I’ve seen teams of five engineers spend six months building a deployment pipeline that nobody outside their group can use without a three-hour onboarding call. Why? Because the team prioritized speed over usability. They hardcoded environment variables, skipped the README, and never considered that another team might have a different directory structure. When the tool inevitably broke during an incident, the original authors were on vacation, and the on-call engineer had to reverse-engineer a shell script at 2 a.m. That’s not a tool; that’s a liability.
This happens because we treat tooling like a one-off project instead of a product. A product—even an internal one—has users, requirements, and a lifecycle. It needs design, testing, documentation, and iteration. When we skip those steps, we’re not saving time; we’re pushing the cost downstream to every engineer who has to use or fix the thing.
The Ownership Vacuum
One of the biggest problems is that dev tools often fall into an ownership gray zone. Product teams own features. Platform teams own infrastructure. But who owns the script that generates test data? Or the local development environment that half the company depends on? Usually, it’s whoever built it last—until they leave or get reassigned.
I once worked on a team where the primary build tool was a massive Makefile written by a senior engineer three years prior. He’d left the company. The Makefile had ballooned to over 2,000 lines, with no comments and a dependency graph that nobody understood. When builds started failing intermittently, we spent two weeks debugging it—two weeks that should have gone into feature work. The eventual fix was a 20-line change, but it took a dozen engineers to understand the problem well enough to make it safely.
This is what happens when tools don’t have dedicated maintainers. They rot. They accumulate technical debt faster than any user-facing system because there’s no product manager arguing for refactoring time, no QA team catching regressions, and no customer complaining when things break—until suddenly, everything breaks at once.

The Real Cost of Neglected Tooling
When I talk about “cost,” I’m not being abstract. I mean concrete, measurable drains on engineering capacity. Let’s break it down:
1. Onboarding becomes a nightmare. A new hire’s first weeks set the tone for their entire tenure. If they spend those weeks wrestling with a broken local setup, outdated wikis, and tribal knowledge passed around in Slack DMs, they’re not learning the codebase—they’re learning survival skills. I’ve seen startups where the onboarding process literally involves a senior engineer sitting next to the new person for two days, manually editing config files because the setup script doesn’t work on anything but the author’s machine.
2. Incident response slows to a crawl. When production goes down, every second counts. If your debugging tool requires a specific version of Python, a VPN connection, and three environment variables that nobody documented, you’re burning precious minutes on tooling instead of the actual problem. I’ve been on calls where we spent more time getting access to logs than fixing the bug those logs revealed.
3. Context switching kills focus. Engineers who have to maintain their own tools—or worse, work around them—lose hours a week to what I call “tool tax.” Fixing a flaky test runner, updating a hardcoded API key, hunting down the right version of a dependency. These aren’t deep work; they’re interruptions that fracture attention and make it impossible to get into a flow state.
4. Good people leave. This is the one that should keep engineering leaders up at night. Talented engineers don’t stick around at companies where they spend 30% of their time fighting their own toolchain. They know they can go somewhere that has a smooth CI pipeline, a one-command dev environment, and internal tools that actually work. Losing a senior engineer over something as fixable as bad tooling is the most expensive unforced error a company can make.
Why “Just Good Enough” Is a Trap
There’s a seductive logic that says internal tools don’t need to be polished because the audience is small and technical. I’ve heard managers say, “Our engineers are smart; they can figure it out.” That’s not pragmatism—it’s arrogance dressed up as efficiency.
The “figure it out” approach assumes that every engineer has the same context, the same mental model, and the same tolerance for friction. They don’t. A new team member or someone from a different squad might have zero context on why a tool works the way it does. When you force them to reverse-engineer your decisions, you’re wasting their cognitive load on trivia instead of the actual problem they’re trying to solve.
I remember a tool we built for generating synthetic traffic to a staging environment. It worked fine for the team that built it, but when another team tried to use it, they spent days debugging because the tool assumed a specific data format that wasn’t documented. The fix was trivial—add a flag to specify the format—but the time lost wasn’t. And that time came directly out of the sprint they were supposed to be delivering.
What a Better Approach Looks Like
Treating dev tools as first-class products doesn’t mean over-engineering them. It means applying the same discipline we use for customer-facing work: clear requirements, design reviews, testing, documentation, and a plan for maintenance. It’s not about perfection; it’s about intentionality.
Here’s what I’ve seen work in practice:
Assign clear ownership. Every tool that multiple teams depend on needs a named owner—not a team, a person. That person is responsible for keeping the tool running, answering questions, and prioritizing improvements. If they leave, ownership transfers explicitly, not by assumption. This sounds bureaucratic, but it’s the difference between a tool that gets updated and one that becomes archaeological.
Write a README that’s actually useful. I don’t mean a paragraph and a shrug. A good README tells you what the tool does, why it exists, how to set it up, common failure modes, and where to get help. It should be the first place someone goes when things break, not the last. If you can’t explain your tool in writing, you probably don’t understand it well enough yourself.
Invest in automated testing for the tool itself. If your build script doesn’t have tests, it will break in ways you won’t notice until production is on fire. Unit tests, integration tests, smoke tests—apply the same rigor you’d use for any other code. A tool without tests is a ticking time bomb.
Design for the person who knows nothing. Assume the user has never seen your tool before, doesn’t know your team’s conventions, and is already frustrated because something else is broken. Clear error messages, sensible defaults, and a --help flag that actually helps—these are not luxuries. They’re basic respect for your colleagues’ time.
Budget for maintenance. When you ship a tool, you’re making an ongoing commitment. That means allocating time in every sprint for fixes, updates, and user support. If your planning process doesn’t account for this, you’re planning to fail. I’ve seen teams adopt a “20% time” model specifically for internal tooling maintenance, and the payoff in reduced firefighting more than justified the investment.

The Incentive Problem
All of this requires changing how we reward engineering work. If promotions and bonuses are tied exclusively to product features, then tooling work will always be a side hustle. I’ve seen brilliant engineers get passed over for promotion because they spent a quarter fixing the CI pipeline instead of shipping a flashy new feature. The pipeline they fixed saved the company hundreds of hours over the next year, but that impact wasn’t visible to the promotion committee because it didn’t map to a product metric.
This is a leadership failure. Engineering managers need to advocate for the value of tooling work, both in performance reviews and in roadmap planning. They need to track and communicate the impact—hours saved, incidents prevented, onboarding time reduced—in terms that the business understands. If you can say, “This tool eliminated 15 hours of manual setup per new hire, and we hired 20 people this year,” that’s a number that gets attention.
It’s also worth recognizing that some of the most impactful engineering work in history has been tooling. The developers who built Git, Docker, or Kubernetes weren’t working on customer-facing features. They were solving problems that made every other engineer more productive. The companies that understand this—the ones that invest in developer experience as a core competency—tend to ship faster and retain talent better. It’s not a coincidence.
What I Do Differently
On my own teams, I’ve adopted a few rules that help keep tooling from becoming a second-class concern:
- Tooling requests go through the same triage as feature requests. If someone needs a new script or a change to an existing tool, it gets a ticket, a priority, and an owner. No more hallway conversations that result in a half-baked shell script dropped in a shared directory.
- Every tool has a “retirement plan.” When we build something, we agree on the conditions under which it gets deprecated or replaced. Tools that don’t have an exit strategy tend to hang around forever, accumulating cruft and confusing people.
- We dogfood our own tools relentlessly. If it’s painful for us, it’ll be painful for everyone else. I’ve made it a practice to do a fresh install of our dev environment on a clean machine every few months, just to see where the friction points are.
- We celebrate tooling wins publicly. When someone ships a tool that saves the team time, we talk about it in standup, in retros, in all-hands. Visibility creates accountability, and accountability creates quality.
It’s Not About Being Nice—It’s About Being Effective
I’m not making a moral argument here. I’m making a practical one. Treating dev tools as second-class citizens is a tax on your own productivity. It slows down every engineer, increases the risk of incidents, and drives away the people you most want to keep. Fixing it doesn’t require a massive reorganization or a new executive role. It requires a shift in mindset: from treating tools as disposable scripts to treating them as products that deserve the same care as anything else we build.
The next time someone on your team says, “Let’s just hack something together for now,” ask them what “now” means. Because in my experience, “now” usually lasts about three years—long enough for the original author to leave, the documentation to rot, and the hack to become a critical dependency that nobody understands. A little bit of upfront discipline saves a whole lot of downstream pain.
We have the skills to build great tools. We just need the organizational will to value that work properly. Until we do, we’ll keep wasting time on avoidable problems, burning out our best people, and wondering why shipping feels so hard.
Frequently Asked Questions
Why do internal dev tools so often lack documentation?
Because documentation is usually the first thing sacrificed when teams are under time pressure. The tool’s author understands it well enough to skip writing things down, and by the time someone else needs it, the author has moved on. It’s also a symptom of the “second-class citizen” mindset—if you don’t view the tool as a real product, you won’t treat its documentation as a real requirement.
How can I convince my manager to invest more in tooling?
Stop talking about it in engineering terms and start talking about it in business terms. Track the hours your team loses to broken tools, flaky environments, or manual setup. Translate that into dollars or feature delays. Most managers respond to data, not complaints. If you can show that a week of tooling work will save a month of engineering time over the next quarter, you’ll get their attention.
What’s the one thing a team can do today to improve their tooling posture?
Pick the single most painful tool or script your team uses—the one that causes the most friction, the most confusion, the most late-night debugging sessions—and spend a day writing a proper README for it. Document what it does, how to set it up, common errors, and where to get help. It won’t fix the code, but it’ll immediately reduce the support burden and make the tool less of a black box. After that, schedule time to add tests and refactor the worst parts.
Is there a risk of over-engineering internal tools?
Yes, and it’s a trap worth avoiding. The goal isn’t to build a beautiful, general-purpose framework that solves every hypothetical problem. It’s to solve the actual problems your team faces today, with the understanding that you’ll iterate as needs change. If you find yourself adding plugin systems or configuration UIs before you’ve validated that the core functionality works, step back. Ship the simplest thing that does the job well, document it, and improve it based on real feedback, not imagined futures.