EXP-006: Model Comparison Test

Experiment ID

EXP-006

Category

AI Behavior Research

About This Research

This research documents how different AI models responded to identical prompts under comparable testing conditions. The experiment focuses on observable differences in explanation style, instruction compliance, and practical reasoning rather than determining which model is “better.” All observations are based solely on the documented responses collected during this experiment.

Research Question

How do different AI models respond to identical prompts under comparable testing conditions?

Objective

Compare the responses of different AI models using identical prompts to observe differences in explanation style, instruction following, practical reasoning, and response structure.

Research Metadata

FieldValue
Experiment IDEXP-006
CategoryAI Behavior Research
StatusCompleted
Date20-08-2026
AI Models TestedChatGPT, Gemini Pro, and Claude.
Number of Models3
Number of Test Prompts3

Test Environment

ItemDescription
TopicAI Model Comparison
Models TestedChatGPT, Gemini Pro, and Claude, using the model versions available during the documented test.
MethodIdentical prompts submitted separately to each model
EvaluationObservational comparison only
ScopeExplanation quality, instruction following, and practical reasoning

Evaluation Criteria

Instruction compliance: Whether the response followed explicit requirements such as the requested word limit, bullet count, or format.

Explanation style: How the response explained the requested concept and how clearly the information was organized.

Practical reasoning: What recommendation the model gave and what conditions, trade-offs, or reasoning it provided to support that recommendation.

Prompts Used

Prompt 1 – Explanation Task

Explain the difference between AI tools and AI models in under 120 words. Include one practical example.

Prompt 2 – Instruction Following Task

Summarize the following in exactly three bullet points.

Artificial intelligence systems can process large amounts of information, identify patterns, generate predictions, and assist humans in making decisions across many industries.

Prompt 3 – Practical Reasoning Task

A company wants to use AI to summarize customer support emails.

Should they use an AI model directly or an AI tool? Explain your reasoning in about 100 words.

Results

TestChatGPTGeminiClaude
Explanation TaskDistinguished AI tools from AI models and included a practical example.Distinguished AI tools from AI models and included a practical example.Distinguished AI tools from AI models, remained within the requested word limit, and provided a practical example with a clear conceptual explanation.
Instruction FollowingProduced exactly three bullet points while closely following the original wording.Produced exactly three bullet points using greater paraphrasing while preserving the original meaning.Produced exactly three bullet points while preserving the key ideas of the original passage with concise wording and minimal paraphrasing.
Practical ReasoningRecommended using an AI tool for most organizations while noting situations where direct model use may be appropriate.Presented a conditional recommendation based on organizational requirements and technical capabilities.Recommended using an AI tool while explaining the engineering overhead of direct model use and identifying scenarios where direct model integration may be appropriate.

Key Observation

Across the documented tests, all three AI models completed the three assigned tasks. The main observable differences were in response style, level of detail, and how recommendations or explanations were presented.

Research Finding

Within the three documented tasks, ChatGPT’s responses were comparatively concise and direct, Gemini’s responses were more detailed and conditional, while Claude’s responses placed greater emphasis on conceptual clarity and practical implementation considerations. These observations describe the responses collected in this experiment and should not be treated as general characteristics of the models.

Limitations

This experiment evaluated three AI models using three documented prompts. The observations are limited to the tested model versions, prompts, and conversation settings. Different prompts, future model updates, or alternative testing conditions may produce different results.

AI Models Tested: ChatGPT (OpenAI Web Tier), Gemini Pro (Google Web Tier), and Claude (Anthropic Web Tier) — Tested under public web interfaces.

Repeatability

This experiment was conducted as a single documented comparison using the workflow described in the Research Methodology. Repeating the experiment with different prompts, model versions, conversation settings, or testing conditions may produce different results.

Why This Matters

AI users often assume different AI models will produce identical responses to the same prompt. This experiment demonstrates that even when models respond to the same task, they can differ in explanation style, level of detail, instruction following, and practical reasoning.

Evidence

Figure 1

ChatGPT response to the explanation task comparing AI tools and AI models using the documented prompt.

ChatGPT explaining the difference between AI tools and AI models with a practical example.
Figure 1. ChatGPT response to the explanation task distinguishing AI tools from AI models.

Figure 2

Gemini response to the explanation task comparing AI tools and AI models using the documented prompt.

Gemini explaining the difference between AI tools and AI models with a practical example.
Figure 2. Gemini response to the explanation task distinguishing AI tools from AI models.

Figure 3

Claude response to the explanation task comparing AI tools and AI models using the documented prompt.

Claude response explaining the difference between AI tools and AI models with a practical example.
Figure 3. Claude response to the explanation task distinguishing AI tools from AI models.

Figure 4

ChatGPT response to the instruction-following task, producing exactly three bullet points from the provided text.

ChatGPT summarizing text into exactly three bullet points as instructed.
Figure 4. ChatGPT response to the instruction-following task using exactly three bullet points.

Figure 5

Gemini response to the instruction-following task, producing exactly three bullet points from the provided text.

Gemini summarizing text into exactly three bullet points while following the given instruction.
Figure 5. Gemini response to the instruction-following task using exactly three bullet points.

Figure 6

Claude response to the instruction-following task, producing exactly three bullet points from the provided text.

Claude response summarizing a passage into exactly three bullet points.
Figure 6. Claude response to the instruction-following task using exactly three bullet points.

Figure 7

ChatGPT response to the practical reasoning task recommending an AI solution for customer support email summarization.

ChatGPT recommending an AI tool for summarizing customer support emails with supporting reasoning.
Figure 7. ChatGPT response to the practical reasoning task recommending an AI solution for customer support email summarization.

Figure 8

Gemini response to the practical reasoning task evaluating AI tool and AI model usage for customer support email summarization.

Gemini evaluating whether an AI tool or AI model is appropriate for summarizing customer support emails.
Figure 8. Gemini response to the practical reasoning task evaluating AI tool and AI model usage for customer support email summarization.

Figure 9

Claude response to the practical reasoning task recommending an AI tool while explaining when direct AI model integration may be appropriate.

Claude response recommending an AI tool for customer support email summarization while explaining when direct AI model use may be appropriate.
Figure 9. Claude response to the practical reasoning task recommending an AI tool while explaining scenarios where direct AI model integration may be appropriate.

Related Articles

Citation

AI Tools Usage Guide Project. (2026). EXP-006: Model Comparison Test (AI Behavior Research Log). Independent AI Behavior Research Series.

Publication Information

Published:
30 July 2026

Last Updated:
20 August 2026

Version:
1.0

Editorial Review:
Completed

Editorial Note

This research log documents a controlled comparison of multiple AI models responding to identical prompts under comparable testing conditions. The observations presented are limited to the documented prompts, model versions, and evaluation criteria used during this experiment and should not be interpreted as definitive rankings or universal measures of model performance.

The purpose of this publication is to support transparent AI behavior research through detailed documentation of the testing process and observations. Differences observed in response style, instruction compliance, and practical reasoning may vary with future model updates, alternative prompts, or different testing environments. Readers should interpret the findings within the methodology and limitations described in this report.

Previous Experiment
EXP-005: Confidence vs. Accuracy Test