,

Study Finds AI Models Struggle With Industrial Parts Questions

Why This Matters to Distributors: Artificial intelligence is moving into product search, customer service and inside sales, but new research suggests general-purpose AI models remain unreliable when questions require exact part numbers, compatibility rules, and product configurations. Web search improved performance in the study, but the highest-scoring configuration still produced 42 incorrect answers out of 100.

Leading artificial intelligence models struggled to answer industrial product questions accurately in a new study that highlights the risks distributors face when deploying general-purpose AI for product selection and technical support.

Reshape Automation’s Industrial AI Accuracy Index 2026 tested GPT-6 Astra, Claude, and Gemini on 100 questions covering products from 14 industrial manufacturers. The questions included part searches, product validation, cross-references, configurations, and technical support.

With web search enabled, GPT-6 Astra recorded the highest score at 52%, followed by Claude at 40.5%. Without search, scores fell sharply, ranging from 12% to 14%. The study gave half credit for partially correct responses.

The results point to a significant gap between the ability of AI models to locate information and their ability to reliably apply detailed industrial product data.

GPT-6 Astra with web search gave 46 fully correct answers, 12 partially correct answers and 42 incorrect answers. Claude with search produced 35 correct, 11 partial and 54 incorrect responses. Across all five configurations tested, 349 of 500 responses were rated incorrect.

Web search made a substantial difference. It increased GPT-6 Astra’s overall score by 40 percentage points and Claude’s by 26.5 points. On conventional part searches, GPT-6 Astra’s score increased by 51.1 points with search, while Claude improved by 35.2 points.

The improvement dropped sharply when models had to interpret product configuration rules rather than retrieve information published on a webpage.

GPT-6 Astra with search scored 22.7% on 11 configuration questions, while Claude with search scored 0%. The report cautioned that 10 of the 11 questions involved one manufacturer’s catalog, limiting how broadly the results can be applied.

The models performed better on straightforward product searches. GPT-6 Astra with search scored 60.2% on 44 part-search questions, compared with 47.7% for Claude. GPT-6 Astra scored 80% on five diagnostics and technical-support questions, although the small sample makes that result less conclusive.

Accuracy was not the only issue identified in the research. The models sometimes produced different answers when asked the same question multiple times.

Each question was submitted three times to each configuration. In 66 of the 500 question-and-configuration combinations, or 13.2%, the repeated responses received different verdicts. Claude with search was inconsistent on 21% of questions, compared with 9% for GPT-6 Astra with search.

In one example, Claude with web search was asked whether a Siemens 3RW5950-0CH00 communications module was compatible with a 3RW52 soft starter. The model answered yes in all three attempts. The study’s verified answer was no because the module was compatible with the 3RW55 series.

The study also found that some incorrect answers were delivered without signaling uncertainty. Based on a model-assisted review of a sample of responses, Reshape estimated that 36.5% of the 349 incorrect answers provided a fact, figure or product code without a hedge, refusal, or request for clarification.

“Distributors are where these questions land,” Reshape Automation CEO and co-founder Juan Aparicio said. “A customer asks whether a part fits, someone on the inside sales team has a few minutes to answer, and one wrong digit ships the wrong part.”

The research has important limitations. Reshape developed the questions, generated the reference answers through its own ReshapeX system, and conducted the evaluations. No independent third party reviewed the study. ReshapeX was not evaluated under the same conditions as the other systems because its verified responses served as the reference answers used to grade them.

Gemini also was not tested with web search, and the study used a Flash-tier Gemini model. All tests were single-turn interactions, and several category and manufacturer samples were small. Reshape disclosed 13 limitations in the report.

Despite those qualifications, the results underscore a practical issue for distributors evaluating AI for technical product applications: models that perform well on general information tasks may still require structured, verified product data and additional controls before their answers can be relied on for part selection, compatibility and configuration.

Do not miss any content from Distribution Strategy Group. Join our list.


Share this article: