· NERVICO · artificial-intelligence · 15 min read
My AI adoption journey: lessons from Mitchell Hashimoto
The HashiCorp founder shares his 6-step roadmap for adopting AI agents: from initial frustration to 20% automation. Harness engineering, slam dunk tasks, and practical lessons for your team.
Mitchell Hashimoto, founder of HashiCorp and creator of tools like Vagrant and Terraform, publicly shares his experience adopting AI agents: a six-month journey from initial frustration to running agents in the background during 10-20% of his workday.
This isn’t theory. It isn’t marketing. These are practical lessons from a senior engineer who went from rejecting inefficient chatbots to systematically integrating agents into his daily workflow, creating a concept he calls “harness engineering” to make agents truly reliable.
In this article you’ll discover his step-by-step methodology, what “slam dunk tasks” he identified as perfect for agents, and how to apply these lessons to your team without falling for the hype. Companies that adopt agents with clear methodology report average returns of 171%, with 74% achieving ROI in the first year. But success isn’t in the technology—it’s in how you adopt it.
Mitchell Hashimoto’s 6-step journey
Step 1: Drop the chatbot
Mitchell’s first discovery was blunt: traditional chatbots (ChatGPT-style) are inefficient for real development work. The copy-paste code interaction, back to editor, test, back to chat, adjust prompt, repeat… is too much friction.
His recommendation: use agents from the start. Agents aren’t glorified chatbots. They’re systems that can execute actions in a loop: read files, run commands, make HTTP requests, see results, and adjust their next action based on what they observe.
Tools like Claude Code, Cursor with agent mode, or GitHub Copilot with agentic capabilities fall into this category. The key difference: the agent lives in your development environment, not in a separate browser tab.
Step 2: Reproduce your own work
This was the phase of greatest friction for Mitchell. Rather than abandon when initial results didn’t meet expectations, he forced himself to reproduce all his manual commits using agents. This meant doubling the work: doing things manually first, then trying to get the agent to reach the same result.
Sounds inefficient. It is. But it revealed three fundamental principles that transformed his use of agents:
1. Break sessions into clear, separate tasks
Agents work best with specific goals. “Implement authentication” is vague. “Create an Express middleware that verifies JWT in the Authorization header and returns 401 if invalid” is specific.
2. Separate planning from execution
Don’t ask the agent to plan and execute simultaneously. First ask for a plan. Review it. Adjust if needed. Then ask it to execute that specific plan. Two distinct phases.
3. Provide verification mechanisms
The agent needs ways to validate its own work. Automated tests. Verification scripts. Commands it can run to confirm it did things right. Without this, the agent advances without knowing if it’s on the right track.
Step 3: End-of-day agents
Mitchell started reserving 30 minutes at the end of each day to launch agents on low-pressure tasks. No urgent deadlines. No need for it to work perfectly. Space to experiment and learn.
Three types of tasks worked especially well:
Deep research on libraries and tools: “Research the 5 most popular libraries for WebSocket handling in Go, compare their trade-offs, and give me basic code examples for each.”
Parallel exploration of speculative ideas: “How would we implement hot-reload in Ghostty without restarting the process? Explore 3 different approaches and their implications.”
GitHub issue triage with CLI tools: “Review the last 20 open issues, classify them by type (bug/feature/docs), identify duplicates, and suggest which should be prioritized.”
The common pattern: useful but not critical tasks. If the agent fails, there are no serious consequences. If it succeeds, you gained valuable time and knowledge.
Step 4: Outsource the “slam dunks”
This is where Mitchell started seeing real net efficiency. A “slam dunk task” is a high-confidence task where the agent will almost certainly succeed. They’re not obvious at first. You identify them through experience.
Examples of his slam dunks:
- Update documentation after API changes
- Implement unit tests for pure functions with clear specifications
- Refactor code to follow project-specific conventions
- Create UI components based on existing designs with slight variations
The key: run these agents in the background while you work on creative or architectural tasks. And here comes a critical point many people ignore: turn off notifications.
When the agent finishes and notifies you, you context-switch. You review its work. It distracts you from what you were doing. Context-switching kills productivity. Better: let the agent work, and review its results at a time you choose, not when it finishes.
Step 5: Harness engineering
This is the most important concept of Mitchell’s journey. “Harness engineering” means systematically preventing agent errors, not just reacting when they fail.
When an agent makes a mistake, Mitchell doesn’t just fix it and move on. He takes time to engineer a solution that prevents that error from happening again. Two main mechanisms:
AGENTS.md files: Documentation that captures discovered bad behaviors. Each line in this file is based on a real error the agent made.
Example from Ghostty’s AGENTS.md (his macOS terminal project):
- NEVER run `make test` without filters. Use `make test FILTER=<pattern>` for specific tests.
- For screenshots, always use the `./scripts/screenshot.sh` script we configured, not manual commands.
- The CI workflow requires all tests to pass. Run `make check` before committing.
- Don't modify files in `vendor/` directly. Use `make vendor-update` instead.Programmed tools: Custom scripts that make it easy for the agent to do things right. If the agent failed taking screenshots manually, Mitchell created screenshot.sh that encapsulates the correct way. If it needed to run filtered tests, he created helpers that make it obvious how to do it.
The result: according to Mitchell, this “almost completely resolved” recurring problems.
The philosophy behind it: “I don’t want to run agents for the sake of running agents.” Agents must provide value, not create more supervision work.
Want to implement AI agents in your team without the chaos? At NERVICO we help you with a free audit of your processes to identify your “slam dunk tasks”.
Step 6: Always have an agent running
Mitchell’s current goal: maintain constant background agent activity. Not multiple agents in parallel (unnecessary complexity), but a single agent working continuously on the next task in the queue.
For this he uses slower but more thoughtful models, like Amp’s “deep” mode, which prioritizes reasoning over speed. The result: during approximately 10-20% of his normal workday, there’s an active agent executing tasks.
This wasn’t a “big bang” adoption. It was iterative, gradual, based on learning from mistakes and building trust step by step.
Harness Engineering: making agents reliable
The harness engineering concept Mitchell created is gaining traction in the industry. Anthropic recently published a technical article on effective harnesses for long-running agents. Philipp Schmid wrote about the importance of agent harness in 2026. It’s not just Mitchell’s idea—it’s an emerging necessity.
What is an Agent Harness
An agent harness is the infrastructure that wraps an AI model to manage long-running tasks. Think of it as the agent’s operating system. Its responsibilities include:
- Memory and state management: The agent needs to remember what it’s done, what it’s learned, what failed.
- Controlled tool integration: Not giving it access to everything, but to the specific tools it needs with clear guardrails.
- Dynamic context engineering: Deciding what information to include in each model request to maximize effectiveness without exploding token limits.
- Planning and task decomposition guidance: Helping the agent break complex problems into manageable steps.
- Verification and safety constraints: Preventing the agent from doing dangerous or costly things without supervision.
Why harnesses matter
Model benchmarks show small differences. Claude Opus 4.6 outperforms GPT-5.3 on certain tests by 2-3%. Seems marginal. But these benchmarks evaluate short, isolated tasks.
In real production, what matters is durability: how well the model follows instructions while executing hundreds of tool calls over hours. A 1% difference on a benchmark doesn’t detect if the model drifts from the goal after 50 steps.
The harness is what keeps the agent on the right path. As the agent engineering community says: “Harness engineering is becoming as important as model selection.”
Harness Engineering vs Prompt Engineering
Prompt engineering optimizes inputs. Harness engineering builds safety and recovery systems around the model.
The analogy: The harness is like a safety harness in climbing. It doesn’t prevent you from attempting difficult routes. It prevents you from falling when you fail. It lets you take calculated risks because you know there’s a safety net.
In practice, Mitchell updates his AGENTS.md every time the agent makes a mistake. It’s not just passive documentation—it’s an active behavior contract that evolves with the team’s experience.
”Slam Dunk Tasks”: where agents really shine
Not all tasks are equal. Mitchell identifies “slam dunks” as high-confidence tasks where the agent will almost certainly succeed. What characterizes them?
Characteristics of slam dunk tasks
- High repeatability: You’ve done this or something very similar many times before.
- Well-defined context: The agent has access to all necessary information.
- Clear success criteria: You can verify if it worked without ambiguity.
- Low ambiguity: Doesn’t require subjective judgment or complex trade-off decisions.
- Automatic verification possible: Tests, linters, or scripts that confirm correctness.
Important: these tasks aren’t obvious at first. You discover them by experimenting. What’s a slam dunk for your team may not be for another.
Real examples
From Mitchell’s experience:
- Deep library research (gather info, compare trade-offs)
- GitHub issue triage (classify, identify duplicates, label)
- UI component recreation with clear specifications
- Junior/mid-level bug fixes with specific context
Mitchell’s critical note: “Changes always need thorough review, follow-ups, and sometimes manual tweaks to reach Senior quality.” Agents aren’t autonomous developers. They’re force multipliers.
Where agents still struggle
Mitchell is honest about limitations:
- Ambiguous requirements: “Make the UI look better” doesn’t work.
- Novel problems without precedent: If it’s never been done before, the agent has no patterns to learn from.
- High-level architecture: Deciding between microservices vs monolith requires business context the agent doesn’t have.
- Business decisions with complex trade-offs: Speed vs security, cost vs features, etc.
- Genuine creativity: Agents combine existing patterns exceptionally well, but creating something truly new remains human.
This honesty is refreshing. It contrasts with the “agents will replace developers” hype we see in marketing.
Industry data confirms the pragmatic approach: Gartner projects that in 2026, 40% of enterprise applications will include task-specific agents. Specific. Not general. The emphasis is on practical high-confidence tasks with clear ROI, not AGI ambitions.
How to bring this to your team (without the chaos)
Mitchell’s journey was individual. How do you translate it to a development team? Here’s a roadmap adapted from his experience combined with enterprise adoption data:
Weeks 1-2: Individual adoption
Goal: Familiarization without pressure.
Each developer experiments with code assistants: Cursor, GitHub Copilot, Claude Code. No productivity expectations. Just exploration.
Expected result: Identify early adopters. There are always 2-3 people on each team who embrace new tools quickly. Let them lead by example, not by imposition.
Weeks 3-4: Team documentation
Goal: Tangible value without risk.
Use agents to update technical docs, READMEs, onboarding guides. These are low-criticality but high-utility tasks. If the agent makes a mistake, it’s easy to correct. If it succeeds, you saved hours of tedious work.
Expected result: First shared win. The team sees concrete value.
Month 2-3: Workflow automation
Goal: Identify and automate team slam dunks.
Start collective harness engineering. Create a team AGENTS.md. Document discovered bad behaviors. Build shared tools.
Candidate tasks: preliminary automated code reviews, bug triage, basic regression testing, dependency updates with validation.
Expected result: 10-15% of routine tasks delegated to agents.
Month 4+: Continuous optimization
Goal: Refine and expand.
Measure what works. Discard what doesn’t. Expand to new tasks gradually. Not from adoption pressure, but from ROI evidence.
Expected result: 20-30% of routine tasks automated, allowing the team to focus on complex and creative problems.
Real success metrics
Don’t trust just feelings. Measure. Companies adopting AI agents with discipline are seeing quantifiable results:
- Average ROI: 171%, with US companies reaching 192%
- Time to ROI: 74% achieve return in the first year
- Productivity: 39% of executives report at least doubled productivity
- Content creation: 46% faster in marketing teams
- Customer service: 52% reduction in time for complex cases (ServiceNow case)
- General productivity: 15-30% gains in areas like customer service
The critical factor according to Oracle’s analysis: “Successful organizations aren’t those with the best AI infrastructure, but those that activate agents quickly, measure tangible results, and iterate with discipline.”
Managing resistance
Adopting agents generates legitimate fears:
“Will I become obsolete?”
Reframe: Agents are multipliers, not replacements. They eliminate tedious work so you can focus on interesting problems. Developers who use agents effectively have an advantage over those who don’t.
“Will juniors learn the fundamentals?”
Mitchell mentions this explicitly. His recommendation: Juniors should learn fundamentals without agents first. Understand why code works, not just how to make an agent generate it. After that, agents are powerful tools.
“Will we lose quality control?”
Only if you don’t implement harness engineering. Review is still necessary. Agents propose, humans decide.
Transparency helps. Share wins and fails. Show AGENTS.md growing with the team’s learning. Make the process visible, not just the results.
What Mitchell learned (so you don’t have to repeat it)
Key lessons
1. Knowing when NOT to use agents is as valuable as knowing when to use them
Don’t use agents for the sake of using agents. Mitchell insists on this: “I don’t want to run agents for the sake of running agents.” Have judgment. Sometimes doing something manually is faster and more effective.
2. Net efficiency came only at step 4
The first three steps were investment, not gain. There was more friction, more work, more frustration. Requires patience and commitment to the process. Most people quit too soon.
3. Notifications off are critical
Context-switching is productivity’s silent enemy. When an agent notifies you it finished, it interrupts your flow. Let agents work in the background. Review their results when you choose.
4. Single agent > Multiple parallel agents
Mitchell runs one agent at a time. Less complexity. More control. Lower API costs. More predictable results. Multi-agent orchestration is tempting in theory, chaotic in practice.
Common mistakes
1. Abandoning in the inefficiency phase
The adoption curve has a valley. At first you’re less productive with agents than without them. It’s temporary. Mitchell forced himself through this friction. Most people don’t.
2. Not documenting errors
Without harness engineering, you repeat the same problems. The agent makes error X. You fix it. A week later, it makes error X again. AGENTS.md is your organizational memory. Use it.
3. Unrealistic expectations
Agents aren’t autonomous senior developers. They need review. They need context. They need guidance. They’re junior-to-mid level at best. If you expect more, you’ll be disappointed.
4. Scaling without validating
Don’t evangelize agents to the whole company because they worked on your team. Validate with specific slam dunks. Measure results. Then expand. Data before enthusiasm.
Mitchell’s disclaimer
Mitchell ends his article with a paragraph worth quoting in full:
“This blog post was fully written by hand, in my own words. I hold no financial stakes in AI companies. I respect individual choices regarding AI adoption. I’m a software craftsman that just wants to build stuff for the love of the game.”
Why it matters: Even early adopters maintain critical perspective. It’s not evangelism. It’s pragmatism. Mitchell acknowledges that in months, his current views may seem naive given how fast the field evolves. Intellectual humility.
Your action plan for the next 2 weeks
Enough theory. How do you start tomorrow?
Week 1: Personal experiment
Day 1-2: Choose a tool. Claude Code, Cursor, or GitHub Copilot. Configure the environment. Read basic documentation. Don’t expect productivity yet.
Day 3-4: Reproduce a manual commit with the agent. Choose something you already did. Try to get the agent to reach the same result. Document what worked and what didn’t.
Day 5: Reserve 30 minutes of “end-of-day agent time”. Launch the agent on low-criticality tasks: research, exploration, documentation. No pressure.
Week 2: First slam dunks
Day 1-2: Identify 3 candidate slam dunk tasks in your workflow. Criteria: repetitive, well-defined, low risk if they fail. Write them specifically.
Day 3-4: Run agent in background on one of those tasks. Notifications off. Work on something else. Review result later.
Day 5: Evaluate honestly. Did it work? If yes, add it to your slam dunk list. If not, update your personal AGENTS.md with what you learned. What failed. Why. How to prevent it.
Repeat this cycle. Gradually your slam dunk list grows. Your AGENTS.md becomes richer. Your confidence in agents increases based on evidence, not hype.
Resources to go deeper
Primary sources:
Technical resources:
- Anthropic: Effective Harnesses for Long-Running Agents
- The Importance of Agent Harness in 2026
- Notes on Agentic Engineering with Mitchell Hashimoto
Industry data:
Conclusion
Mitchell Hashimoto didn’t start believing in AI agents. He went through frustration, inefficiency, and many mistakes before reaching 10-20% of his workday with agents active in the background. His journey reveals that successful adoption isn’t about technology—it’s about methodology.
Harness engineering, slam dunk tasks, and patience to push through initial friction phases. These are the keys. Not the newest model. Not the most hyped tool. Disciplined process and iterative learning.
Companies adopting agents with this discipline (measurement, iteration, clear ROI) are seeing 171% returns. Those doing it for hype are wasting resources. The difference isn’t the technology—it’s the approach.
You don’t need revolution. You need evolution. Start tomorrow with a 2-week personal experiment. Document what you learn. Build your harness. Identify your slam dunks. In three months, you’ll have measurable results.
As Mitchell says: “I’m a software craftsman that just wants to build stuff for the love of the game.” AI agents are tools. Good tools, if used well. Not magic. Not replacements. Capacity multipliers in the hands of professionals who know what they’re doing.
Ready to start? Discover how we implement technical teams with AI agents using our APEX framework, or request a free audit to identify opportunities in your current workflow.