Experiment ID
EXP-006
Category
AI Behavior Research
About This Research
This research documents how different AI models responded to identical prompts under comparable testing conditions. The experiment focuses on observable differences in explanation style, instruction compliance, and practical reasoning rather than determining which model is “better.” All observations are based solely on the documented responses collected during this experiment.
Research Question
How do different AI models respond to identical prompts under comparable testing conditions?
Objective
Compare the responses of different AI models using identical prompts to observe differences in explanation style, instruction following, practical reasoning, and response structure.
Research Metadata
| Field | Value |
|---|---|
| Experiment ID | EXP-006 |
| Category | AI Behavior Research |
| Status | Completed |
| Date | 20-08-2026 |
| AI Models Tested | ChatGPT, Gemini Pro, and Claude. |
| Number of Models | 3 |
| Number of Test Prompts | 3 |
Test Environment
| Item | Description |
|---|---|
| Topic | AI Model Comparison |
| Models Tested | ChatGPT, Gemini Pro, and Claude, using the model versions available during the documented test. |
| Method | Identical prompts submitted separately to each model |
| Evaluation | Observational comparison only |
| Scope | Explanation quality, instruction following, and practical reasoning |
Evaluation Criteria
Instruction compliance: Whether the response followed explicit requirements such as the requested word limit, bullet count, or format.
Explanation style: How the response explained the requested concept and how clearly the information was organized.
Practical reasoning: What recommendation the model gave and what conditions, trade-offs, or reasoning it provided to support that recommendation.
Prompts Used
Prompt 1 – Explanation Task
Explain the difference between AI tools and AI models in under 120 words. Include one practical example.
Prompt 2 – Instruction Following Task
Summarize the following in exactly three bullet points.
Artificial intelligence systems can process large amounts of information, identify patterns, generate predictions, and assist humans in making decisions across many industries.
Prompt 3 – Practical Reasoning Task
A company wants to use AI to summarize customer support emails.
Should they use an AI model directly or an AI tool? Explain your reasoning in about 100 words.
Results
| Test | ChatGPT | Gemini | Claude |
|---|
| Explanation Task | Distinguished AI tools from AI models and included a practical example. | Distinguished AI tools from AI models and included a practical example. | Distinguished AI tools from AI models, remained within the requested word limit, and provided a practical example with a clear conceptual explanation. |
| Instruction Following | Produced exactly three bullet points while closely following the original wording. | Produced exactly three bullet points using greater paraphrasing while preserving the original meaning. | Produced exactly three bullet points while preserving the key ideas of the original passage with concise wording and minimal paraphrasing. |
| Practical Reasoning | Recommended using an AI tool for most organizations while noting situations where direct model use may be appropriate. | Presented a conditional recommendation based on organizational requirements and technical capabilities. | Recommended using an AI tool while explaining the engineering overhead of direct model use and identifying scenarios where direct model integration may be appropriate. |
Key Observation
Across the documented tests, all three AI models completed the three assigned tasks. The main observable differences were in response style, level of detail, and how recommendations or explanations were presented.
Research Finding
Within the three documented tasks, ChatGPT’s responses were comparatively concise and direct, Gemini’s responses were more detailed and conditional, while Claude’s responses placed greater emphasis on conceptual clarity and practical implementation considerations. These observations describe the responses collected in this experiment and should not be treated as general characteristics of the models.
Limitations
This experiment evaluated three AI models using three documented prompts. The observations are limited to the tested model versions, prompts, and conversation settings. Different prompts, future model updates, or alternative testing conditions may produce different results.
AI Models Tested: ChatGPT (OpenAI Web Tier), Gemini Pro (Google Web Tier), and Claude (Anthropic Web Tier) — Tested under public web interfaces.
Repeatability
This experiment was conducted as a single documented comparison using the workflow described in the Research Methodology. Repeating the experiment with different prompts, model versions, conversation settings, or testing conditions may produce different results.
Why This Matters
AI users often assume different AI models will produce identical responses to the same prompt. This experiment demonstrates that even when models respond to the same task, they can differ in explanation style, level of detail, instruction following, and practical reasoning.
Evidence
Figure 1
ChatGPT response to the explanation task comparing AI tools and AI models using the documented prompt.

Figure 2
Gemini response to the explanation task comparing AI tools and AI models using the documented prompt.

Figure 3
Claude response to the explanation task comparing AI tools and AI models using the documented prompt.

Figure 4
ChatGPT response to the instruction-following task, producing exactly three bullet points from the provided text.

Figure 5
Gemini response to the instruction-following task, producing exactly three bullet points from the provided text.

Figure 6
Claude response to the instruction-following task, producing exactly three bullet points from the provided text.

Figure 7
ChatGPT response to the practical reasoning task recommending an AI solution for customer support email summarization.

Figure 8
Gemini response to the practical reasoning task evaluating AI tool and AI model usage for customer support email summarization.

Figure 9
Claude response to the practical reasoning task recommending an AI tool while explaining when direct AI model integration may be appropriate.

Related Articles
- AI Tools vs. AI Models: Why ChatGPT, Gemini, and Claude Give Different Answers
- How Prompt Structure Controls AI Output (The Logic Test)
- How to Choose the Right AI Tool: A Workflow-Based Evaluation
- What Is an AI Workflow? Beginner Guide With a Real 3-Step Example
- Why AI Tools Behave Unpredictably Compared to Traditional Software
Citation
AI Tools Usage Guide Project. (2026). EXP-006: Model Comparison Test (AI Behavior Research Log). Independent AI Behavior Research Series.
Publication Information
Published:
30 July 2026
Last Updated:
20 August 2026
Version:
1.0
Editorial Review:
Completed
Editorial Note
This research log documents a controlled comparison of multiple AI models responding to identical prompts under comparable testing conditions. The observations presented are limited to the documented prompts, model versions, and evaluation criteria used during this experiment and should not be interpreted as definitive rankings or universal measures of model performance.
The purpose of this publication is to support transparent AI behavior research through detailed documentation of the testing process and observations. Differences observed in response style, instruction compliance, and practical reasoning may vary with future model updates, alternative prompts, or different testing environments. Readers should interpret the findings within the methodology and limitations described in this report.
Previous Experiment
EXP-005: Confidence vs. Accuracy Test