What do B2B buyers ask ChatGPT when researching software and comparing vendors? Surveys and prompt research offer insights into how buyers use AI to explore options, evaluate solutions, and make decisions. But the exact questions B2B buyers ask, and how often they ask them, remain unclear. This article examines what existing evidence reveals and where the gaps remain. It also explains how B2B marketers can uncover the questions their own buyers ask to make their AI search strategy more relevant.
Three kinds of prompt data, and only one is typed by a buyer
Most claims about "what buyers ask" rest on one of three things, and they answer different questions.
- Logs are records of what real users typed. OpenAI's usage research and donated chat exports are logs. They show wording and frequency, but rarely who the user was or whether they were buying anything.
- Surveys are what people say they do. A buyer survey can tell you that comparing vendors is the top reason people open a chatbot during software research. It cannot tell you how the question was worded, and recall is not behavior.
- Tracked lists are the prompts a team chose to monitor. They reflect the team's judgment about what buyers might ask. A list of 50 tracked prompts says what the team decided to check, not what anyone typed.
Mixing the three is where most errors start. A tracked list gets described as "what buyers ask." A survey percentage gets read as a share of real prompts. A general-user log gets treated as B2B behavior. Keep the source type attached to every claim, and the evidence below sorts itself out.
What the largest usage research can and cannot say about buying
The biggest public log study is the NBER working paper How People Use ChatGPT, published in September 2025. OpenAI's economic research team ran it with outside academics. It classified roughly 1.1 million sampled conversations from May 2024 to June 2025, drawn from consumer plans only, using automated classifiers and no human reading of message content.
Its main finding for our question is the shape of the questions. By the paper's main analysis, about 49 percent of messages were asking, 40 percent were doing, and 11 percent were expressing. Seeking information grew from 14 percent to 24 percent of usage over the year to mid-2025. And by June 2025 about 73 percent of messages were non-work.
The paper's taxonomy does include a sub-category for our topic. "Purchasable Products" sits under seeking information and is defined as inquiries about products or services available for purchase, with an example like asking for a laptop recommendation under $1,000. It is a defined category. The text we could read does not report its share, and the sample covers consumer plans, not business accounts. So the paper tells us people use ChatGPT to ask more than to do, and gives us a label for product questions, but it does not tell us how often a B2B buyer asks one.
Our piece on what independent data shows about commercial intent in ChatGPT covers that gap in more detail.
A second public dataset shows how much the sample matters. A late-2025 study of 142,808 publicly shared conversations across five chatbots found seeking information at about 39.6 percent of requests, averaged across platforms. That is a different unit, taxonomy and sample from OpenAI's 24 percent, and people choose what to share. Neither number is wrong. They answer different questions, which is the point of the section above.
What buyers say they ask
Surveys are the only evidence that speaks to B2B buyers directly. Here is what each one says about questions and tasks, and what to hold back.
| Source | Sample and date | What it says about questions and tasks | Caveat |
|---|---|---|---|
| G2, Answer Economy Report | 1,076 B2B decision-makers, March 2026 | Over two-thirds start with a category or competitor question; about one in five start with questions about requirements or process. Comparing vendor strengths and weaknesses is the top reason for using chatbots in software research | G2 runs a software review marketplace and has an interest in how buyers research software. Self-reported; the base for the first-prompt question is not stated |
| Software Finder survey, as reported by MarketingProfs | 714 B2B software buyers, March 2026 | 47% use AI tools frequently when researching software vendors; about a third use AI to generate side-by-side comparisons during evaluation; 68% say AI introduced them to vendors they would not have found otherwise; 18% say AI is the most influential source for building a shortlist | Secondary reporting of the survey; self-reported |
| Gartner, press release | 645 B2B buyers, August to September 2025, released May 2026 | 45% used generative AI in a recent purchase, mainly to gather information on vendors and products | Press release from the party that ran the survey; question wording not given |
| Responsive survey, as reported by Digital Commerce 360 | More than 350 B2B buyers; dates not stated | Buyers use conversational tools to summarize, compare and recommend vendors; they typically start with five to eight vendors and narrow to three or fewer | Secondary reporting; the publisher sells response-management software |
Read together, the surveys agree on the main jobs: find vendors, compare them, and gather product information. They do not agree on how to count, and none of them reports the words buyers used.
One reading deserves care. If over two-thirds of buyers start with a category question or a competitor question, then most first prompts do not contain your brand name. Our inference: a tracked list built mostly from brand-name prompts, such as "[us] pricing" or "[us] reviews", watches the later stages of a buyer's questions and misses the opening. G2's report does not define how it sorted prompts into categories, so treat the split as a direction, not a measurement.
What prompts look like when someone can see them
Researchers can sometimes see real prompts. A September 2026 preprint from the Max Planck Institute for Software Systems and partner universities, Characterizing Web Search by Conversational LLM Agents, analyzed chat exports that volunteers donated through data-access requests. The dataset holds 171,264 conversations from 613 users across ChatGPT, Claude, Grok and DeepSeek, with ChatGPT the largest share.

The prompts were long. The authors report that nearly 80 percent of user prompts contain more than 20 terms. The searches the engines sent were short: almost all generated web queries had fewer than 10 to 15 terms, and the authors cite 2 to 4 terms as typical for a human search query. On ChatGPT turns where the queries were visible, the average was 3.07 web queries per prompt, issued in parallel. Our piece on query fan-out explains that mechanism.
Two more details matter for anyone building a prompt list. First, engines do not search on every turn. By our arithmetic from the paper's Table 1, ChatGPT used web search on about 6 percent of turns in the donated sample, and the authors report that the share of search-triggering prompts rose over time on every platform. Second, the sample is volunteers recruited on a crowd-work platform, covering general use, not B2B buying. It tells us about the form of prompts, not their subject.
The form is the useful part. A buyer who types a paragraph with constraints is asking a different question from one who types two keywords, and the answers differ. Our piece on how query phrasing changes which B2B brands get cited covers that effect.
Seven jobs, and what backs each one
The G2 report names the jobs its respondents give chatbots in software research. Its text lists seven: comparing vendor strengths and weaknesses, basic product research, vendor identification, use case validation, drafting requests for proposals, working through pricing and packaging, and validating fit. It ranks only the first. The chart sorts them by how much independent support each has.

Every job is named in one survey. Two have support from a second survey: comparing vendors (about a third of Software Finder's respondents generate side-by-side comparisons) and vendor identification (68 percent say AI introduced them to vendors they would not have found otherwise). Basic product research has partial support from Gartner's finding that buyers mainly use generative AI to gather information on vendors and products. The last column is empty for all seven: none of the sources we found measures how often each job is typed.
The jobs also span the whole purchase. Identification and basic research sit early. Comparison, use case validation and fit sit in the middle. Drafting RFP questions and working through pricing sit late, close to a buying decision. A prompt list that covers only "best of" and pricing leaves out most of that range.
What this changes about a B2B prompt list
The evidence is thin on frequencies and fairly consistent on form and range. Five common habits and what to do instead:
| Common habit | What the evidence suggests | Do this instead |
|---|---|---|
| Keyword-style prompts such as "best AP software" | Real prompts run past 20 terms and carry constraints | Write full sentences with company size, stack, region and constraints a buyer would state |
| Mostly brand-name prompts | Surveys say most first prompts start from a category or a competitor | Include unbranded category prompts and competitor-led prompts, not only your own name |
| One prompt per topic | One prompt becomes several searches, and wording changes answers | Write two or three phrasings per question and keep them fixed across checks |
| Pricing and "best of" only | Buyers name seven jobs, from identification to fit | Cover all seven jobs, including use case validation, fit and RFP-style questions |
| A list built in a brainstorm | A tracked list shows what a team decided to check | Source the wording from calls, tickets and other places buyers speak |
None of these changes depends on knowing the exact share of each question type. They depend only on the form and range of buyer questions, which the evidence supports better than it supports any frequency. Long, specific prompts are also the subject of our piece on long-tail B2B AI visibility, and a list of short head terms misses them.
Where your own buyers' wording lives
The best source for what your buyers ask is the places they already speak to you. Each one skews in a known direction.
| Source | What it shows | What it skews toward |
|---|---|---|
| Sales call recordings and notes | Questions in the buyer's own words, with context | Late-stage buyers who already know you exist |
| Lost-deal and win/loss notes | Why a buyer chose someone else, and what they compared | Deals that reached a decision; misses early researchers |
| Support tickets and onboarding questions | What customers needed to know to use the product | Existing customers, not prospects |
| Incoming RFPs and security questionnaires | Formal requirements and the order buyers list them in | Larger buyers and regulated categories |
| On-site search and chat transcripts | What visitors looked for after they arrived | Visitors who already found your site |
| Search Console queries of ten or more words | Long queries that already lead to your pages | Buyers who found you through Google |
| Community threads where buyers ask peers | Unprompted questions and the language peers use | Public, English-language communities and vocal users |
Teams that already keep call and ticket research in one place, for example in a customer and market research tool such as Omnibound's Intelligent Research, can pull this faster. A spreadsheet works too. Collect at least ten verbatim questions for each job before you write a single prompt, record where each came from, and keep the original wording next to the cleaned version.
From a sales call quote to a trackable prompt
Verbatim buyer language is messy. The steps are to strip names, keep the constraints, make each entry one question, and write it as a full sentence. Three invented examples show the change.
| What the buyer said (invented) | Cleaned question | Trackable prompt |
|---|---|---|
| "We're on NetSuite and our AP team is three people, so nothing that needs a consultant." | AP automation for a small team on NetSuite without a long implementation | Which accounts payable automation tools work with NetSuite for a three-person team that cannot take on a long implementation? |
| "The last vendor had a platform fee we didn't see coming." | Extra fees beyond the per-invoice price | What fees do accounts payable automation vendors usually charge beyond the per-invoice price, such as platform, onboarding or integration fees? |
| "Our auditors want an approval trail." | Audit trail for approvals | Which accounts payable automation tools keep a full approval trail that an external auditor can review? |
Each prompt now holds the constraint the buyer stated. Tag each with its job (identification, pricing, validation) and whether it names a brand, so coverage can be counted by job later.
Questions people ask
What is the most common first question a B2B buyer asks a chatbot?
Surveys point to a category question or a competitor question. G2 reports that over two-thirds of its respondents start that way. No log-based source reports the share for B2B buyers.
Do B2B buyers actually use ChatGPT to research vendors?
Surveys say many do. Gartner reports 45 percent used generative AI in a recent purchase, and Software Finder reports 47 percent use AI tools frequently when researching software vendors. Both are self-reported. Our piece on commercial intent in ChatGPT sets the survey figures side by side.
Are chatbot prompts different from search queries?
In the donated-log study, yes: nearly 80 percent of prompts ran past 20 terms, against the 2 to 4 terms the authors cite for a typical search query. The sample is general use, not B2B.
Is it enough to track prompts that name my brand?
The surveys suggest not. If most first prompts start from a category or a competitor, brand-name prompts watch the later stages and miss the opening. Our reading, not a measurement.
How many prompts should an inventory hold?
No source gives a number. A workable rule is enough entries per job that no prompt in your tracked list rests on a single wording, and a refresh whenever sales hears a new question.
From a question inventory to a prompt list you can track
An inventory is the first draft of a prompt list. Once you have ten or so entries per job, run a handful through the AI Search Visibility Checker to see whether your brand shows up for questions written the way a buyer would put them. Then use AI Search Intelligence to track each prompt individually over time, so you can see which jobs you appear in, which you miss, and whether that changes after you publish or earn a new mention.
Turn Your Content Into AI-Search Winners
Get cited across ChatGPT, Claude & Perplexity — not just ranked on Google.
- Increase AI citations
- Improve answer visibility
- Track brand mentions in LLMs