skills (Claude Code)
Claude Code skills for the parts of software work that land under your name: reviewing someone else's pull request, answering the reviews on your own, and taking a tracked bug all the way to a merge-ready pull request. Each one is a written method with hard gates, installable as a plugin marketplace.
Why it exists
A coding agent is at its most dangerous when its output leaves the machine. A review comment posted on a colleague’s pull request, a reply that tells a reviewer they were wrong, a bug “fixed” by a patch that only silences the symptom: each lands under the user’s name, and each is hard to take back.
Left to themselves, agents fail here in predictable ways. They summarise a diff instead of reviewing it. They report findings whose evidence is “this looks like it might”. They agree with whichever reviewer sounds most confident, bots included, and change correct code to please them. They declare a bug fixed because a test went green, without checking that the test could ever fail.
oxn holds an agent to a project’s architecture. This repository holds it to an engineering process: a written method for each of these jobs, with gates the agent has to pass before anything goes out.
How it works
Each skill is a SKILL.md that Claude Code loads when the task matches, or when it is invoked by slash command. bugfix is the exception: it pushes code and opens pull requests, so it runs only when typed as a command. There are three:
pr-reviewreviews someone else’s pull request. It checks the PR out into its own git worktree, so the user’s checkout is never touched, and gives the test suite a scratch database so an unmerged migration never reaches their dev data. It runs the suite at the merge base and at the PR head, and only the difference counts as a finding. Two subagents sweep the diff for breadth while the main agent reads it in full for depth.pr-feedbackworks the reviews on your own pull request. Every comment, from a human or a bot, is treated as a claim and gets exactly one verdict (Confirmed, Confirmed-partly, Unconfirmed, Contradicted, Preference or Contested), each with its own evidence bar. When two reviewers want opposite things, both sides are verified and handed to the user; the agent does not pick.bugfixtakes a tracker issue (Jira, Linear, GitHub Issues) from report to pull request: reproduce it on the latest default branch as a failing test, state the root cause, have a fresh context try to refute it, fix it, open the PR, get CI green, and answer the review bots.
All three are installable together as a Claude Code plugin marketplace, or one at a time by copying a directory.
Design decisions
- Evidence or silence. A finding without a
file:linethat makes it true is dropped. A reproduced failure outranks an argument. A bot’s confidence is not evidence, and neither is a reviewer’s seniority. - Print, ask, then post. Everything outward-facing is printed to the terminal first, and nothing is posted until the user says which parts go out. Threads the agent disagreed with stay open: the reviewer decides whether the answer settles it.
- A test has to be able to fail.
bugfixcommits the fix on its own, reverts it with the new test still in place, and requires the test to fail again with the same failure signature it had before the fix. A test that fails for a different reason is coupled to the fix, not to the bug, and the gate rejects it. Existing tests are read-only: the agent may not loosen one to get to green. - Fresh eyes over self-review. A model re-reading its own reasoning does not catch its own mistakes, so diagnoses and diffs go to a subagent that sees only the evidence. Optional flags bring in second opinions from other model families, and they skip silently when the tool is not installed.
- A method, not a toolchain. The skills never hard-code a test or lint command. They find the repository’s real ones, CI workflows first, and say so when a step has no equivalent.
- Ask at named points only. Each skill lists the few places it may stop to ask. Everywhere else it decides, so a run does not turn into a stream of confirmations.
Status
Actively developed: three skills so far, with more as I turn other parts of my workflow into them. MIT licensed.
Click any figure or table to zoom.