Questions to Ask a Codex Agency Before You Hire

Questions to Ask a Codex Agency Before You Hire

Founder of Goodspeed

The questions you ask a Codex agency before hiring them are the cheapest and most powerful diligence you have. A good agency will welcome hard questions, because they have real answers. A weak one will deflect, generalise, or steer you back to the vision. The quality of the answers, and the ease with which they are given, tells you almost everything you need to know.

We are an AI engineering team that ships production software with agents like Codex and Claude Code, and we sit on the receiving end of these questions regularly. The best clients ask sharp ones. This piece gives you the questions that actually separate serious agencies from the rest, along with what a strong answer sounds like and what should worry you.

Use these in your first conversations, and pay as much attention to how they answer as to what they say.

Can you show me production software you have built?

This is the first and most important question. You want to see live software with real users, not demos, not decks, and not prototypes. A demo proves an agency can make something look good in a controlled setting. Production software proves they can build things that survive contact with reality, which is the only thing that matters once you have paid.

A strong answer is immediate and specific: here are systems that are live, here is what they do, here is who uses them. Listen for detail about what happened after launch, because that is where the real engineering lives. A weak answer leans on demos, hypotheticals, or vague references to work they cannot show you. If an agency cannot point to software running in the real world, treat everything else they say with caution.

How do you use Codex and agents in your workflow?

This question reveals whether they understand the tool or just wield it. You are listening for a grounded, specific account of where agents genuinely help and where human judgement takes over. Strong practitioners have a clear mental model: they know which tasks to hand to an agent, which to keep by hand, and why. They talk about the agent as a powerful tool with limits, not a magic box.

A weak answer treats Codex as the achievement itself, as if using it were the value. It is not. The value is the software that results. Be wary of anyone who sells the tool rather than the outcome, or who cannot explain their own process for deciding what the agent does. The way an agency talks about their workflow tells you whether they are engineers using a tool or enthusiasts riding a trend.

How do you stop an agent shipping something wrong?

This is the question that separates real engineering teams from the rest. Agents generate code quickly, which means they generate mistakes quickly, and without guardrails those mistakes ship. A serious agency will answer this readily, because they have been burned before and built the scaffolding to prevent it: automated tests, evaluation harnesses, type safety, code review, staged rollouts, and monitoring.

Listen for a culture where a human owns every line that goes live. The best teams do not trust agent output by default; they verify it. A weak answer waves the question away, implies the model is reliable enough not to need checking, or has no real process behind it. If nobody is rigorously reviewing what the agent produces, you are not hiring an engineering team. You are hiring a fast way to accumulate hidden problems.

Who will actually work on my project?

Some agencies sell you senior expertise and then hand the work to whoever is free, with the agent papering over the gaps. Because agent output can look uniform regardless of who prompted it, this is easier to hide than it used to be. You have a right to know who will build your software, what their experience is, and how the work is reviewed.

A strong answer names the people, describes their experience, and explains the review process that keeps quality consistent. Ideally you meet the people who will actually build, not just the person who sells. A weak answer is vague about staffing or evasive about seniority. The presence of the agent does not remove the need for experienced engineers; it raises it, because someone has to supply the judgement the agent lacks. Know whose judgement you are buying.

Who owns the code, accounts, and infrastructure?

Ownership is where agent-built projects quietly go wrong, so ask early and directly. You want to own the code, the accounts, the infrastructure, and the credentials outright, with nothing locked behind the agency's private tooling. Some setups leave you unable to touch your own software without going back to the people who built it, which is a weak position you do not want to discover late.

A strong answer is plain and comfortable: you own everything, you get full access, and there are no hidden dependencies. Good agencies have nothing to hide here and will say so without hesitation. A weak answer is vague, hedged, or reluctant, and that hesitation is one of the most useful warning signs you can pick up. Software you cannot fully control is a liability, and this question surfaces the risk while it is still cheap to walk away.

What does handover look like if we part ways?

Closely related to ownership is what happens when the relationship ends. You want to know that you can take the software and continue without the agency, whether that means running it yourself or moving to another team. This tests whether they are building something you genuinely own or something that keeps you dependent on them.

A strong answer describes a clean handover: full access, documentation, and a deliberate transfer of knowledge so your people or another team can pick it up. It signals confidence and integrity, because an agency comfortable with you leaving is one that expects to keep you by being good, not by trapping you. A weak answer treats handover as an awkward afterthought. How an agency talks about the end of the relationship tells you a lot about how they will behave during it.

How do you handle security and sensitive data?

Every serious build touches security, and agent-built software especially, because agents can introduce weaknesses as quickly as they introduce features. Ask how the agency thinks about security, how they protect sensitive data, and how they make sure agent-generated code does not open holes. The answer reveals whether security is built into their process or bolted on at the end.

A strong answer treats security as a first-class concern with real practices behind it, and can speak to the specific standards your situation requires. If you operate under compliance obligations, ask directly whether they have built software that met them. A weak answer is thin or generic, which is a genuine problem if your software handles anything sensitive. The depth of the security answer scales with the stakes, so weight it according to what your software will actually hold.

How do you support software after launch?

Launch is the beginning, not the end. Software changes, requirements shift, dependencies age, and AI-built systems in particular need ongoing attention as models and tools evolve. Ask what post-launch support looks like: who fixes things when they break, how quickly, on what terms, and how the software is monitored once it is live.

A strong answer treats maintenance as a core part of the work, with a clear model for support, monitoring, and ongoing improvement. It shows they build for the long term, not just to hit a launch date and move on. A weak answer goes quiet on everything after go-live, which sets you up for a cliff exactly when you start depending on the software. How an agency talks about the period after launch tells you whether they build things meant to last.

How do you handle scope, milestones, and change?

Agent-accelerated projects can drift, because it is easy to keep generating more without ever finishing anything properly. Ask how they define scope, set milestones, and handle changes along the way. You want to understand how you will know the work is progressing and how you will agree what done means for each piece.

A strong answer describes clear scope, explicit milestones, and a sensible process for handling change without letting the project sprawl. Good agencies push towards this because it protects them as much as you. A weak answer is loose about how work is defined and tracked, which with fast tooling tends to produce an ever-growing pile of half-features and no working whole. Clarity here is cheap insurance, and its absence is a reliable predictor of a messy engagement.

What happens when something goes wrong?

Every real project hits problems. What matters is how the agency handles them. Ask what happens when a build runs into trouble, a deadline is at risk, or something breaks in production. The answer tells you how they behave under pressure, which is when you most need them to be honest and capable rather than defensive.

A strong answer is candid and specific: here is how we communicate bad news, here is how we triage and fix, here is how we keep you informed. Teams that have shipped real software have real stories of things going wrong and being put right, and they tell them without flinching. A weak answer implies nothing ever goes wrong, which is either dishonest or inexperienced. You are hiring for the hard days as much as the easy ones.

A note on certifications and how to weigh the answers

Somewhere in your diligence, an agency may imply an official Codex qualification. Be clear with yourself that there is no formal Codex certification, so any official-sounding credential is either a misunderstanding or embellishment. Assess genuine expertise instead, which is production evidence, clear engineering practice, honest answers about limitations, and software that has survived real use.

Taken together, these questions are less about collecting facts and more about reading behaviour. Weight ease and specificity heavily. The agencies worth hiring answer hard questions comfortably and in detail, because they have lived the answers. The ones to avoid deflect, generalise, or retreat to the vision. If you only remember one thing, remember to watch how they answer as closely as what they say, and to trust production evidence above every claim.

Conclusion

The right questions turn a sales conversation into real diligence. Ask to see production software, how they use agents, how they stop bad code shipping, who does the work, who owns it, how handover and security and maintenance are handled, and how they behave when things go wrong. Then weight the answers by how comfortably and specifically they are given, and ignore any certification claim, because none exists.

Do this and you will quickly separate the agencies that ship reliable software from the ones selling a story. The answers, and the ease of them, will tell you almost everything.

If you want a team that ships production software with AI coding agents, see our AI work, or book a free call with our AI engineering team.

Harish Malhi - founder of Goodspeed

Written By

Founder of Goodspeed