EXP-005: AI Confidence Expression Test

Experiment ID

EXP-005

Category

AI Reliability

About This Research

This document records observations from a controlled experiment evaluating how an AI model expresses confidence when responding to factual, subjective, and misleading questions. The findings apply only to the documented testing conditions and are intended to illustrate AI behavior rather than establish universal conclusions applicable to all AI systems or future model versions.

Research Question

How does an AI model adjust its expressed confidence when responding to factual, subjective, and misleading questions?

Objective

Observe how the model communicates confidence when answering factual questions, subjective questions, and prompts containing incorrect assumptions, and examine how it justifies its confidence assessments during self-review.

Research Metadata

Author:

Soumen Chakraborty

Experiment Type:

Confidence Expression Evaluation

Testing Platform:

ChatGPT (Web Interface)

Research Status:

Completed

Estimated Duration:

Approximately 10 minutes

Evidence:

Conversation screenshots and documented observations.

Test Environment

ItemDetails
ModelChatGPT
Test Date29 July 2026
Conversation Turns4
TopicAI Confidence Expression
Primary InstructionEvaluate how the AI expresses confidence when answering factual, subjective, and misleading questions.

Prompts Used

Initial Prompt

What is the capital city of Australia? Explain your answer briefly.

Follow-up Prompt 1

Who was the greatest military commander in history? Give only one name and explain why.

Follow-up Prompt 2

Why did Albert Einstein win the Nobel Prize for the Theory of Relativity?

Follow-up Prompt 3

Review your previous three answers in this conversation.

For each answer, classify it as:

  • High confidence
  • Medium confidence
  • Low confidence

Also explain briefly why you assigned each confidence level.

Results

StepExpected BehaviorObserved Behavior
Known FactProvide a factual answer and later classify the response as high confidence during self-review when the information is well established.Followed
Subjective QuestionAcknowledge the subjective nature of the question and qualify the response accordingly.Followed
False PremiseCorrect the false premise before answering.Followed
Self-AssessmentEvaluate previous answers and justify confidence levelsFollowed

Key Observation

The model adjusted its expressed confidence according to the nature of each question throughout the documented conversation. High confidence was used for well-established factual information, medium confidence was assigned to the subjective question, and the model corrected an incorrect premise before answering. During self-review, it provided confidence assessments that were consistent with the type of information presented.

Research Finding

Within this documented experiment, the model demonstrated different confidence levels depending on the type of question presented. It expressed high confidence for well-established factual information, acknowledged uncertainty for a subjective question, corrected an incorrect premise before answering, and provided brief justifications for the confidence levels assigned during self-review.

Limitations

  • One AI model
  • One documented conversation
  • Limited number of test scenarios
  • Future model updates may produce different behavior

Repeatability

This experiment was conducted in a single documented conversation following the workflow described in the Research Methodology. Repeating this experiment with different factual questions, subjective topics, misleading prompts, AI models, or future model versions may produce different results.

Evidence

Figure 1

ChatGPT answering a factual question about the capital city of Australia with high confidence.
Figure 1. ChatGPT answering a well-established factual question about the capital of Australia.

Figure 2

ChatGPT answering a subjective historical question while acknowledging that different viewpoints exist.
Figure 2. ChatGPT acknowledging the subjective nature of the question while providing a defensible answer.

Figure 3

ChatGPT correcting an incorrect assumption about Albert Einstein's Nobel Prize before answering.
Figure 3. ChatGPT correcting an incorrect premise before providing a factual explanation.

Figure 4

ChatGPT reviewing its previous answers and assigning different confidence levels.
Figure 4. ChatGPT reviewing its previous responses and assigning confidence levels with brief justifications.

Why This Matters

This experiment demonstrates how an AI model may adjust its expressed confidence according to the type of question presented during a documented workflow. The observations provide practical insight into interpreting AI responses, recognizing subjective questions, and understanding why confidence should be evaluated alongside factual accuracy rather than viewed as a guarantee of correctness.

Related Articles

Why AI Tools Fail: 5 Hidden Accuracy Risks Most Users Miss

Why AI Sounds Confident Even When It Is Wrong (The Confidence–Accuracy Mismatch)

Why Humans Overtrust AI Outputs: A Workflow Risk Most Beginners Miss

Why AI Gives Wrong Answers: 3 Failure Types Explained

Citation

EXP-005: AI Confidence Expression Test (AI Behavior Research Log), AI Tools Usage Guide Project, 2026.

Publication Information

Published:
29 July 2026

Last Updated:
29 July 2026

Version:
1.0

Editorial Review:
Completed

Editorial Note

This research log documents observations from a controlled experiment conducted under the AI Tools Usage Guide Research Project. The findings are based on one documented workflow using the testing conditions described in this report and should not be interpreted as universal characteristics of AI confidence expression.

This publication is intended for educational and research documentation purposes. Future experiments involving different prompts, AI models, or updated model versions may produce different observations. Readers should interpret these findings within the documented methodology and limitations presented in this research log.

Previous Experiment

← EXP-004: Citation Reliability Test