Research

Automotive AI Answer Accuracy: An Open Study Protocol

Sep 7, 2026 7 min read

By SocialCRM · Sources checked September 7, 2026

This is SocialCRM’s proposed method for studying factual accuracy in AI answers about automotive service businesses. It is a research protocol, not a completed study. No shops have been evaluated under this published protocol, and no accuracy rate or improvement result is claimed here.

Research question and unit of analysis

The proposed question is: how accurately does a specified AI system describe the services, location, hours, and qualifications of a defined sample of independent repair shops? The primary unit is one factual claim in one saved answer. Discovery mentions and citations are separate observations.

The protocol deliberately avoids scoring whether a shop is “best.” That term depends on customer needs and supporting evidence. It also avoids claiming consumer exposure from provider API tests. Consumer-interface observations and API observations must be reported in separate groups.

Define the sample before collecting answers

For a pilot, select a bounded group of shops with accessible public facts and an owner or authorized reviewer who can verify disputed information. Record the selection method, locations, specialties, and exclusions before testing. Report the pilot as that sample; do not describe a convenience sample as representative of the aftermarket.

Freeze the question set and review rules before collecting answers. Record any later change as a protocol amendment. A proposed pilot can use three question types per shop and two collection dates. These are design choices to test, not a statistically validated sample size.

Define the sample before collecting answers
Question typeReusable promptWhat to review
Known shop factsWhere is [shop] in [city], and what are its opening hours?Address and hours against dated sources
Service fitDoes [shop] in [city] service [vehicle or service]?Scope and conditions against shop evidence
Local discoveryWhich shops in [city] offer [service]? Include sources.Mentions and citations; verify any factual claims

Build the reference facts first

Create a dated reference record for each shop before testing. Store the public source URL, retrieval date, fact, and evidence excerpt. Ask an authorized shop representative to resolve conflicting facts when feasible. Record when a reference fact remains unknown.

Owner confirmation can help adjudicate a fact, but distinguish that confirmation from publicly accessible evidence. Do not publish private customer records, credentials, repair histories, or personal contact details. If a public source changes during collection, preserve the relevant dated record and flag the change.

Collect enough context to repeat the check

  • Assign a study ID, shop ID, question ID, and run ID.
  • Record the exact prompt, date and time, system, model if exposed, interface type, and available search settings.
  • Record declared location, language, and relevant account or session context without publishing private account information.
  • Save the complete answer and cited URLs where permitted. Retain a permitted excerpt or reference if full publication is restricted.
  • Record timeouts, refusals, missing answers, and inaccessible citations. Do not silently retry until a favorable answer appears.
  • Apply the same retry policy and collection window to every shop. Log each retry separately.

Review each factual claim with explicit labels

Split an answer into claims before labeling it. “The shop is open Sundays and services electric vehicles” contains two claims. Use two reviewers for the pilot when practical. Record their initial decisions, disagreements, and final adjudication separately.

Also record whether a citation actually supports the nearby claim. A working link does not establish support. A missing mention is a discovery observation, not automatically a factual error. A missing answer is a collection outcome, not a zero-accuracy answer.

Review each factual claim with explicit labels
LabelRule
SupportedThe dated reference evidence supports the claim
ContradictedThe dated reference evidence conflicts with the claim
UnverifiedThe available evidence cannot resolve the claim
Not factualThe statement is an opinion, recommendation, or other non-verifiable text

Report results without hiding the denominator

Publish the numbers of attempted checks, completed answers, failed checks, factual claims, supported claims, contradicted claims, and unverified claims. Report each system and interface separately. If you calculate a supported-claim proportion, divide supported claims by all factual claims, including unverified claims, and show the counts.

Claims within one answer and answers about one shop are related. Do not treat every claim as an independent customer observation. A small pilot should emphasize counts and examples with limitations rather than broad industry estimates.

For correction follow-up, record the approved source change and repeat the frozen questions. Report before-and-after observations. Without a design that addresses other changes, do not claim that the edit caused the answer difference.

Reusable collection worksheet

Copy these field names into a spreadsheet or evaluation system. Keep one row per factual claim. Keep a separate run log so failed checks remain visible.

  • Before collection: publish the scope, selection method, questions, retry policy, and planned reporting rules.
  • Before analysis: reconcile reviewer disagreements and preserve unresolved claims.
  • Before publication: check source permissions and remove private information.
  • With results: publish limitations, protocol amendments, a permitted dataset, and a correction contact.
study_id, shop_id, question_id, run_id, collected_at, system, model, interface_type, location_context, prompt, answer_reference, claim_text, citation_url, reference_url, reference_checked_at, initial_review_1, initial_review_2, final_label, adjudication_note, correction_date, follow_up_run_id

Sources and further reading

Platform guidance is linked below. The workflows and study design are SocialCRM recommendations, not ranking guarantees.

Put the guidance to work

Start with a recorded check of what AI says about your shop. Compare the answer with verified facts before deciding what to change.

Start your AI Shop Checkup

Keep reading