Red Flags to Watch For When Hiring a Codex Agency

Red Flags to Watch For When Hiring a Codex Agency

Founder of Goodspeed

Hiring a Codex agency is easy to get wrong, because the market is new and the sales gloss is thick. When a whole industry is racing to attach itself to AI coding agents, plenty of teams learn the vocabulary long before they learn the craft. The danger is that the warning signs are subtle if you do not know what to look for, and expensive if you find out too late.

The good news is that the red flags are consistent. Teams that cannot actually ship reliable production software built with OpenAI Codex tend to give themselves away in the same ways: no real production track record, no evaluations, no guardrails, and a lot of hype standing in for evidence. Once you can name the patterns, you can spot them in a single conversation.

This piece is a field guide to those warning signs. Treat any one of them as a reason to dig deeper, and treat several together as a reason to walk away, no matter how good the pitch sounds.

No production software to point to

The biggest red flag is the absence of software running in the real world. If everything an agency shows you is a demo, a prototype or a concept, they have not proven they can finish. Production is where the hard problems live, and a team that has never shipped into it is asking you to fund their first attempt at senior rates.

Ask directly for systems that are live today, used by real people, handling real data. A confident team names them and explains what they do. A team that dodges, or keeps redirecting to what they could build, is telling you they have not yet built much that lasted. Take that dodge seriously, because the gap between a demo and production is where most projects fail.

They cannot explain how they test

If an agency cannot describe how they verify that their software works, that is a serious warning. Professional engineering rests on testing and, for AI features, evaluations that measure accuracy against known cases. A team that treats it works when I try it as sufficient has no way of knowing whether the system will hold up once real users arrive.

This is one of the easiest flags for a non technical buyer to catch. Ask how they know the software is correct. A vague answer, or visible discomfort, tells you they ship on hope rather than evidence. Without evaluations you have no early warning system, which means you find out about defects when your customers do, at the worst possible moment.

No guardrails around the AI

If the software uses AI at runtime and the team has no answer for what happens when it misbehaves, walk carefully. Guardrails, the controls that validate inputs, filter outputs, cap costs and provide safe fallbacks, are what stop an AI feature causing real damage. Their absence signals a team that has not run AI in anger.

Ask what happens when the model returns something wrong, expensive or harmful. A serious team has concrete mechanisms and can describe them. A team that says the AI is reliable, or looks surprised by the question, has never seen a model fail at scale. That naivety becomes your liability the day the system meets unpredictable real world input.

Hype where evidence should be

Beware the pitch built on adjectives. Revolutionary, cutting edge, game changing and next generation are words, not evidence. When an agency leans on hype about AI rather than showing you specific systems and specific results, the hype is usually there to fill a hole where the track record should be.

Genuine capability sounds different. It is specific and often modest: here is what we built, here is how it works, here is what was hard, here is the result. If you find yourself impressed but unable to point to a single concrete thing the team has actually delivered, step back. The most confident marketing frequently sits on top of the thinnest substance.

Promises that ignore trade offs

Any team that tells you AI can do anything, or that your project carries no real risk, is either inexperienced or dishonest. Real software is a chain of trade offs, and people who have built it know where the compromises and dangers sit. The absence of any caution is itself a warning.

Experienced engineers temper enthusiasm with realism. They will tell you what is hard, what might not work, and what they would need to test before committing. A partner who only ever tells you what you want to hear will keep doing that when the project runs into trouble, right up until the moment it becomes undeniable and unfixable.

One person holds everything

A single operator who is the only person who understands your system is a structural risk, however talented they are. If they take leave, fall ill or move on, your project stalls and your codebase becomes a mystery. Enterprise grade work needs more than one person who can maintain what was built.

Ask who will actually build the system and what happens if that person disappears. A resilient team has shared knowledge and internal review, so no single absence is catastrophic. A team that cannot answer, or waves the question away, is exposing you to a failure mode that has sunk plenty of projects the moment the one indispensable person walked out.

Reluctance to hand over the code

You should own the software you pay for, cleanly and completely. An agency that is cagey about source code ownership, repository access or documentation is signalling that they want to keep you dependent. Lock in is a business model for weak teams, because it means you cannot leave even when you want to.

Raise ownership and handover early. A confident team welcomes it, because their work is built to be handed over and to outlast the engagement. Evasiveness suggests the codebase is either a fragile tangle or deliberately opaque. Either way, a partner who resists giving you full control of your own asset is not a partner you want.

Flaunting certificates that do not exist

There is no official Codex certification, so any team parading one as proof of expertise is trading on your uncertainty. The tools are too new for a meaningful formal credential to exist, and leaning on a badge instead of a portfolio is a tell that the portfolio is thin.

Real expertise is shown, not certified. It lives in live systems, in the ability to explain hard decisions, and in honest accounts of what went wrong. If an agency substitutes a qualification for that evidence, ask to see the evidence instead. The way they respond to that request will usually tell you everything you need to know.

Vague, unverifiable case studies

We helped a leading client achieve great results is not a case study, it is a sentence. When the stories an agency tells have no named system, no describable outcome and no reachable reference, treat them as decoration rather than proof. Real work leaves specifics behind, and specifics are exactly what weak case studies lack.

Push for detail. What was built, who used it, what changed, can you speak to them. A team with genuine wins answers happily and connects you to people who confirm the story. A team that only offers warm generalities, and never a verifiable specific, is hoping you will not ask the follow up questions.

No process, just output

If code goes straight from an AI coding agent into your live system with no review, no testing and no staging, that is not speed, it is recklessness. Professional engineering runs work through version control, code review and continuous integration for good reason. A team that has skipped all of it is shipping unverified code into production and calling it efficiency.

Ask how a change gets from an idea to your users. A mature team describes review gates and staged environments. A team that describes generating code and pushing it live has no safety net, which means every mistake the AI makes reaches your customers directly. That is a flag worth taking very seriously.

Pressure and artificial urgency

Be cautious of any agency that rushes you, manufactures scarcity, or pushes you to commit before you have done your checks. High pressure sales tactics are a classic way to stop a buyer thinking clearly, and confident teams rarely need them because their evidence does the persuading for them.

A partner worth having is comfortable with your diligence. They expect you to ask for proof, take references and compare options, because they know they will come out well. If someone is trying to hurry you past those steps, ask yourself what they would rather you did not have time to discover, and slow down until you have.

When to walk away

No single flag is always fatal, but they cluster. A team with no production track record usually also lacks evals, guardrails and a real process, and compensates with hype and pressure. When you see several of these together, the picture is clear regardless of how polished the pitch is, and the right move is to walk.

The cost of ignoring the signs is not just wasted budget, it is months lost and a system you cannot trust or maintain. Insisting on evidence is not being difficult, it is basic self protection. Trust the pattern of red flags over the quality of the sales performance, and you will avoid the partners who cannot actually deliver.

Conclusion

Red flags rarely arrive alone. A team with no production track record usually also has no evals, no guardrails and a lot of hype to paper over the gaps. The pattern is consistent enough that a single careful conversation will surface most of it, if you ask for evidence and watch how the answers land.

You are not being difficult by insisting on proof. You are protecting yourself from becoming the client who funds someone else's learning curve. Hold the line on evidence, walk away when the flags stack up, and you will end up with a partner who can actually deliver.

If you want a team that ships production software with AI coding agents, see our AI work, or book a free call with our AI engineering team.

Harish Malhi - founder of Goodspeed

Written By

Founder of Goodspeed