I gave ChatGPT, Claude and Gemini the exact same request to quote a piece of custom software, twice each, and got six prices back ranging from $42,000 to $128,000. The part that stuck with me is that the $42,000 quote included more of what the customer asked for than the $98,000 quote did.

I made a video walking through the whole experiment and what I think it means for business owners who are starting to lean on AI. It's below, and this post covers the same ground in writing, with links to the actual chats so you can check my work.

Jump to a section of the video

Why I don't blindly trust it

I have a background as an electrician, and I've asked AI basic electrical questions and gotten wrong answers back with confidence. If I wasn't an electrician I would have followed that advice, and depending on the question, that could mean creating a fire hazard inside someone's house.

That's what bothers me. AI often sounds just as confident when an answer is shaky as when it's solid, so you can't tell how reliable an answer is by how it sounds. I also get people reaching out about a project with a list of what AI told them can be done, and some of it is true at the core, but not all of it.

So I set up a test.

The experiment

I wrote the kind of first email a small service business might actually send. A septic pumping and drain company in central Ohio with three trucks, six field techs and two people in the office, where every job gets written on a paper schedule, retyped into Word for the service report, and then entered again into QuickBooks to invoice. The company is made up, but the request is realistic, and the app they're asking for isn't anything crazy.

I pasted the same prompt into ChatGPT, Claude and Gemini, then did it again in a fresh chat with the same settings. I asked each one for a single total price instead of a range, a timeline in weeks, a line item breakdown, and the assumptions behind the number, and I told it not to ask me any clarifying questions.

Read the exact prompt
You are quoting a custom software project. Here is the inquiry.

We run a septic pumping and drain service in central Ohio. Three trucks, six field
techs, two people in the office. Right now every job starts as a phone call, gets
written on a paper schedule, and the techs text photos back to the office. The
office retypes everything into Word to build the service report, emails it to the
customer, then re-enters it again into QuickBooks to invoice. Reports take about
20 minutes each and we run 8 to 14 jobs a day.

We want a web app the techs can use on their phones in the field to see their
schedule, capture job details and photos, get a signature from the customer on
site, and generate the report automatically. The office needs a calendar to
dispatch, a customer list, and a way to see what has been invoiced. Customers
should be able to request service and pull up their past reports without calling us.

Give me:
1. One total price for the build. A single number, not a range.
2. A timeline in weeks.
3. A line-item breakdown of what that number covers.
4. The assumptions you had to make to produce the number.

Do not ask me clarifying questions. Make your best assumptions and commit to a number.

The six quotes

ToolPriceTimelinePuts invoices into QuickBooks?Works with no cell signal?
Gemini$42,00012 weeksYes, automaticallyYes
ChatGPT$45,00010 weeksYesNot mentioned
Gemini$48,50012 weeksYes, automaticallyYes
ChatGPT$70,00014 weeksYes, plus invoice statusLeft out
Claude$98,00016 weeksNo, excluded in writingNot mentioned
Claude$128,00016 weeksYes, both directionsYes

The highest quote is about three times the lowest. Each tool also disagreed with itself: ChatGPT's two answers were $25,000 apart, Claude's were $30,000 apart, and Gemini's were $6,500 apart.

The two ChatGPT answers surprised me the most, because they reasoned through the problem the same way both times. Both worked out that the office loses somewhere around 2.7 to 4.7 hours a day on reports, both assumed QuickBooks Online, both assumed a web app, and they still landed $25,000 apart.

The cheaper quote did more

Automatic invoicing was the main thing this company asked for, since they described retyping every job into QuickBooks by hand. The $42,000 quote included automatic invoicing and working without cell signal. The $98,000 quote included neither, and it said so in writing: "We are NOT writing invoices into QuickBooks automatically."

If you only asked once, you'd have no way of knowing which of these you got.

You can read the chats yourself: ChatGPT at $45,000, ChatGPT at $70,000, Claude at $98,000 and Claude at $128,000. My Gemini account won't create public links, so both Gemini quotes are shown on screen at 3:43 in the video.

Why the answers are so different

It isn't looking the answer up, because there's no price list inside it. It builds the most plausible sounding answer one piece at a time, and $42,000 and $128,000 are both plausible, because real shops have charged both. My guess is that a lot of what it learned from is what people charged for software before AI changed how it gets built, although some shops still charge that today.

Where the email didn't say anything, it filled in the blanks on its own, and it filled them in differently each time. Nobody mentioned cell signal, and three of the quotes included working offline anyway.

This isn't hallucination, which is when AI states something as fact that isn't true. Nothing was made up, and every one of these quotes could be defended. None of them is exactly wrong, but none of them is right either, and that's the problem with relying on it.

It also tends to agree with you

The second problem has a name, sycophancy, which is a fancy word for sucking up to you. These models are trained partly on people rating answers, and people tend to rate the answers they agree with higher, so agreeing got rewarded.

OpenAI has said as much themselves. In April 2025 they rolled back a ChatGPT update for being too agreeable, and wrote that "GPT-4o skewed towards responses that were overly supportive but disingenuous." It isn't only one company, either. A research paper tested five leading AI assistants and found they "consistently exhibit sycophancy," and the part I wanted to highlight is that human feedback "may also encourage model responses that match user beliefs over truthful ones."

It's a lot like the employee who never says no. They're always saying yes, and you never really know what's going on. If you tell it your idea is good, it will usually find reasons it's good, and if you push back, it will often agree that it made a mistake. It's not going to be the disagreeable voice in your life unless you set it up to be one.

When you're running a business and asking it a lot of questions, it's easy to take the answers to heart instead of with a grain of salt.

How much should you trust it?

The question I'd ask is what happens if this answer is wrong. Going back to the electrical example, if an outlet isn't working and the answer is wrong, someone could get shocked or a fire could start, so that's not one to take lightly. For most business work, it sorts out something like this:

A calculator gives you the same output every time you give it the same input, which is called deterministic. AI chat doesn't work that way out of the box, and as more work gets handed to AI, somebody still has to do the quality check and make sure what it did is correct and safe. We wrote more about where AI fits in a mid-sized company in what AI actually means for a mid-sized company.

Getting more consistent answers

The setting that controls how much answers vary is called temperature. You can turn it down when you use these models through their API, but most chat apps, including the ones I used for this test, don't give you that control. In a regular chat, a few things help:

Being specific helps more than anything else, because the less the AI has to assume, the better the answer gets. If you ask it for ice cream, you'll get a wide range of answers. If you tell it you're in Kansas City, you want a boutique ice cream shop with really good ratings to take your wife to, you don't want just vanilla and chocolate, and you want the top five, you'll get something you can actually use.

The same thing applies when you plug AI into a business process. I worked on a prompt library for one of our clients that a lot of their users rely on, and the only way to get quality, consistent output was to make each prompt extremely specific to one small task. One prompt I built for it has now been used over 50 million times.

Getting more honest answers

At this point you're collecting information, not making the decision.

You don't own the model

The chat app and the model behind it belong to the company that makes it, and they update and retire versions on their own schedule. OpenAI publishes the models it's retiring, and its standard is at least 6 months of notice for generally available models, while preview models "may be retired with much shorter notice, such as 2 weeks."

So if AI sits inside your quoting or your compliance, your process can change without you changing anything. A model that gives you good answers in your industry today can start skewing a different way after an update, and for the same reason, this post will age quickly.

What actually moves the price

When we look at a job like this one, a few questions move the price more than anything on the feature list. Does the app only read from QuickBooks, or does it write to it? Reading is much safer, because it can't change your data, and writing is where things get a little scary. Does it have to work with no cell signal? How many years of old records need to be moved into the new system?

Two of those, whether it writes to QuickBooks and whether it works offline, are blanks the AI filled in differently from one chat to the next.

The takeaway

AI is good at helping you see where the risk is, and it can get you thinking about things you weren't thinking about, which is a real asset. It's not good at telling you what something will cost, and it can tell you compliance rules for your industry that are out of date or don't apply to you at all. Use it to think, not to decide, and anything that needs the same answer every time should be code.

If you have questions about any of this, reach out.

Glossary

Model
The engine behind the chat app.
Sycophancy
The model agreeing with you because agreeing got rewarded.
Temperature
A setting that controls how much answers vary.
Deprecated
Retired by the vendor, on a date they choose.
Deterministic
Same input, same output, every time. A calculator is deterministic. AI chat usually is not.
Hallucination
Stating something as fact that is not true.
Offline support
Keeps working with no cell signal.
Data migration
Moving old records into a new system.

Sources

Frequently asked questions

Can I trust an AI quote for custom software?
Not as a number to budget from. When I gave ChatGPT, Claude and Gemini the same request twice each, the quotes ranged from $42,000 to $128,000, and the $42,000 quote included more of what was asked for than the $98,000 one. It can help you see where the risk and the open questions are, but the price depends on blanks it fills in differently every time.

Why does AI give different answers to the same question?
It isn't looking the answer up. It builds the most plausible sounding answer one piece at a time, and where your question leaves something unsaid, it fills in the blank on its own, often differently each time. Asking for a fixed format, sticking with one model version, and being very specific all make the answers more consistent.

What is AI sycophancy?
It's the tendency of AI models to agree with you. Models are trained partly on people rating answers, and people tend to rate agreeable answers higher. OpenAI rolled back a ChatGPT update in April 2025 for being overly supportive, and a research paper found five leading assistants consistently showed this behavior.

How do I decide how much to trust an AI answer?
Ask what happens if the answer is wrong. Drafts and summaries that a person reads first are fine to lean on. Anything a customer sees or anything that becomes a number in your system needs a person to check it. Pricing rules, compliance checks and anything that has to give the same answer every time should be code, not a chat prompt.