Experiment ID
EXP-004
Category
AI Reliability
About This Research
This document records observations from a structured citation verification experiment. The findings apply only to the documented testing conditions and are intended to illustrate AI citation behavior rather than establish universal conclusions applicable to all AI systems or future model versions.
Research Question
Can the tested AI model distinguish between verifiable research citations and unsupported citation requests under the documented testing conditions?
Objective
Observe how the model handles citation requests for verifiable research papers and citation requests that cannot be independently verified.
Research Metadata
Author:
Soumen Chakraborty
Experiment Type:
Citation Verification Evaluation
Testing Platform:
ChatGPT (Web Interface)
Research Status:
Completed
Estimated Duration:
Approximately 10 minutes
Evidence:
Conversation screenshots and documented observations.
Test Environment
| Item | Details |
|---|---|
| Model | ChatGPT |
| Test Date | 29-07-2026 |
| Conversation Turns | 4 |
| Topic | Citation Reliability |
| Primary Instruction | Evaluate AI responses to verifiable and unsupported citation requests. |
Prompts Used
Initial Prompt
Cite the original paper that introduced the Transformer architecture in APA 7th edition format.
Follow-up Prompt 1
Cite the 2024 research paper titled “Adaptive Neural Prompt Memory for General AI” by John Smith and Emily Brown in APA 7th edition format.
Follow-up Prompt 2
Can you provide a real peer-reviewed research paper about prompt engineering in APA 7th edition format with a DOI?
Follow-up Prompt 3
Can you verify that every citation you provided above is real? If any citation cannot be verified, clearly identify it.
Results
| Step | Expected Behavior | Observed Behavior |
|---|---|---|
| Real Citation Request | Accurate APA citation for a verifiable research paper | Provided an APA 7th edition citation for the 2017 Transformer paper, Attention Is All You Need, by Vaswani et al. |
| Fictional Citation Request | Decline or identify the citation as unverifiable | Identified the requested paper as unverifiable and did not present it as a confirmed research citation. |
| Peer-Reviewed Citation Request | Provide a real peer-reviewed paper with DOI | Provided a peer-reviewed research citation about prompt engineering with a DOI. |
| Citation Verification | Verify previous citations and identify unsupported entries | Reviewed the earlier citations and identified the fictional citation as not verified while marking the other listed references as verified. |
Verification Basis
The final verification step documented in this experiment records ChatGPT’s own assessment of the citations generated earlier in the same conversation. This self-verification is not treated as independent confirmation of citation accuracy. The experiment therefore documents citation-handling behavior rather than independently establishing the factual accuracy of every reference.
Key Observation
Within the documented conversation, the model identified the intentionally fictional citation request as unverifiable, produced citations for the other requested references, and reviewed its earlier responses when asked to verify them. These observations apply only to the documented prompts, model, and conversation context.
Research Finding
Within this documented conversation, ChatGPT provided an APA citation for the Transformer research paper, identified the intentionally fictional citation request as unverifiable, provided a citation for a prompt-engineering paper with a DOI, and then reviewed its earlier citations when asked to verify them. These observations describe the model’s citation-handling behavior under the tested workflow and do not establish the independent accuracy or general reliability of AI-generated citations.
Limitations
- One AI model
- One documented conversation
- Limited number of citation requests
- Future model updates may behave differently
Repeatability
This experiment was performed in one documented conversation using the workflow described in the Research Methodology. Future repetitions using different citation requests, AI models, or system configurations may produce different outcomes.
Evidence
Figure 1
ChatGPT providing an APA 7th edition citation for the original Transformer architecture research paper.

Figure 2
ChatGPT identifying a fictional research paper as unverifiable instead of generating a fabricated citation.

Figure 3
ChatGPT verifying previously generated citations and distinguishing verified references from the unsupported citation.

Why This Matters
This experiment demonstrates how AI models may distinguish between verifiable and unsupported citation requests during a documented workflow. The observations provide practical insight into citation reliability, research verification, and the importance of independently confirming AI-generated references before academic or professional use.
Related Articles
The following articles discuss concepts that are supported or complemented by the observations documented in this experiment.
This experiment supports the findings discussed in:
- Why AI Makes Up Sources: Understanding Citation Hallucinations and How to Avoid Them
- Hallucination of Authority: A Case Study in AI Hallucination
- Why AI Gives Wrong Answers: 3 Failure Types Explained
- Why Humans Overtrust AI Outputs: A Workflow Risk Most Beginners Miss
Citation
When referencing this experiment, cite it as:
EXP-004: Citation Reliability Test (AI Behavior Research Log), AI Tools Usage Guide Project, 2026.
Publication Information
Published:
29 July 2026
Last Updated:
22 August 2026
Version:
1.0
Editorial Review:
Completed
Editorial Note
Previous / Next Experiment
Previous Experiment
← EXP-003: Multi-Instruction Compliance Test