Stop AI from shipping broken code with a 'Refutation Gate' pattern
Job to be done: Prevent AI from shipping broken code through a structured review process
🇳🇬 Ways to use this in Nigeria
Ideas to get you started, adapt to your situation.
- Student
Use Claude to find bugs in Python code ChatGPT wrote for your assignment before submitting it.
- 9-5 employee
Use a second AI to test Python code your first AI wrote for a work task, catching errors before deployment.
What you’ll get
You will learn a proven method to catch errors in AI-generated code before it causes problems. Instead of asking an AI to “review” its own work, you’ll use a second AI to actively try and break the code, forcing it to find hidden flaws. This approach works because it changes the AI’s task from confirming correctness to finding failure, which it is much better at.
Tools you need
- ChatGPT (freemium): An AI chatbot used to generate the initial code.
- Claude (freemium): A different AI chatbot used to review and attempt to break the code.
Steps
-
Generate your initial code: Open your preferred AI chatbot (e.g., ChatGPT) and ask it to write the code you need. Be as specific as possible in your request. The author doesn’t share their exact prompt; a starting point:
Write a Python function that takes a list of numbers and returns their average, handling empty lists by returning 0.You should get a block of code that attempts to fulfill your request.
-
Prepare for the “Refutation Gate” review: Open a different AI chatbot (e.g., Claude) or start a completely new chat session with a different model if your tool allows it. The key is to use a “fresh context” that has no memory of the code being written. Using a different model family (like Claude instead of ChatGPT) is recommended for better results, as different models have different blind spots.
-
Apply the “break-it” brief: Paste the code generated in Step 1 into the new AI chat. Then, immediately follow it with the specific “break-it” prompt. This prompt instructs the AI to find ways the code could fail, rather than just confirming it’s correct.
This code is broken. I know it is — I just don't know how yet. Your job is to produce the specific input, sequence, or state that makes it fail. Assume: - the network drops a packet at the worst possible moment - two of these run at the same time - the database write fails AFTER the external call succeeds - the user does the thing no sane user…You should see the AI analyze the code and suggest specific scenarios or inputs that could cause it to fail, along with explanations. This helps you identify potential bugs before deployment.
Original source
This workflow is based on a blog post by infoinlet1, shared on the DEV Community platform. The author spent 30 days using AI to write 100% of their code to discover effective strategies for preventing AI from confidently generating faulty code.
Notes & variations
- Free-tier viability: Most popular AI chatbots like ChatGPT, Claude, and Gemini offer free tiers that are sufficient for this workflow. You can use one for code generation and another for the “Refutation Gate” review without needing paid subscriptions.
- Common mistake: Do not ask the AI that wrote the code to “review” its own work in the same chat session. This is ineffective because the AI will simply agree with itself, adding a false sense of security. Always use a fresh chat context, ideally with a different AI model.
- Tip for better results: When giving the “break-it” brief, add specific scenarios relevant to your code’s context. For example, if your code handles payments, you might add “Assume: the payment gateway times out after processing the charge but before confirming success.” The more specific you are about potential failure points, the better the AI can find them.