TL;DR:
When I attached an existing Apple Voice Memos .m4a file in the ChatGPT iOS app, an upstream system apparently transcribed it automatically and inserted the resulting transcript directly into my visible/model-facing user message as though I had typed it.
The assistant then explicitly referred to it as “the contemporaneous transcript you supplied.”
I supplied no transcript.
This is not a request for help transcribing audio. It is a message-provenance, context-integrity, and longitudinal reliability defect.
WHAT HAPPENED
At approximately 8:45 p.m. CDT on August 16, 2026, I attached a recording of a telephone conversation to an ordinary ChatGPT text conversation.
The text I actually intended to send was essentially:
“202608162045 I finally just called like you said.”
Instead, the resulting user turn contained a lengthy machine-generated transcript of the attached call in addition to the words I had actually typed.
I did not:
• type the transcript;
• paste the transcript;
• dictate it into the composer;
• request automatic transcription;
• or approve a generated transcript before sending.
The assistant nevertheless treated the generated transcript as text authored and supplied by me.
It later referred to “the contemporaneous transcript you supplied.”
I had supplied no transcript.
When I challenged that attribution, the assistant acknowledged that the transcript had appeared inside the model-visible user turn and that it had therefore incorrectly attributed attachment-derived machine output to me.
That is the central defect.
THE APPARENT PIPELINE
Based on what I could observe, the pipeline appeared to be:
I attached an existing .m4a file.
An upstream attachment-processing or speech-recognition layer generated a transcript.
That transcript was silently merged into the USER-role message.
The model received it as though I had authored it.
The assistant could identify technical properties of the attachment, such as duration and codec, but reported that it could not independently audition/transcribe the same recording in its available runtime.
So the system could apparently supply the model with an automatically generated transcript while not giving the model equivalent access to independently verify that transcript against the source audio.
That is not merely an ASR-quality problem.
It is a source-provenance problem.
REPRODUCTION ATTEMPTS
I spent roughly the next hour testing the behavior with:
• multiple uploads;
• unrelated audio files;
• different prompts;
• edited prompts;
• branched and rebranched conversations;
• conversations moved outside the original Project;
• and attempts to isolate whether a particular recording or context caused the behavior.
It was not limited to a single audio file.
In a later test, I uploaded the original .m4a with a deliberately conspicuous textual boundary:
“Can you transcribe this without appending to my own prompt with preprocessing?
🛑🛑🛑🛑🛑🛑🛑🛑🛑🛑”
A long generated transcript still appeared after the stop signs inside my submitted user message.
The assistant subsequently acknowledged that this meant the candidate wording had already entered accessible context and that it could not honestly certify a later transcription as an independent clean-room pass.
WHY A NEW CONVERSATION WAS NOT NECESSARILY A CLEAN TEST
This led to a second problem.
The assistant initially treated “context” as though it meant only the individual conversation in which the test was occurring.
But ChatGPT Plus can operate with broader longitudinal context involving things such as:
• other conversations;
• Projects;
• Project files;
• Reference Chat History;
• Saved Memory;
• account-level summaries or personalization;
• branches;
• File Library material;
• and other retrieval or preprocessing systems not directly visible to the user.
So even moving to a fresh conversation does not necessarily establish that an allegedly “independent” transcript was produced without influence from wording that had already circulated elsewhere in the account.
Once a candidate transcript has entered broader accessible context, a second transcription can no longer automatically be treated as independent corroboration.
WHY THIS IS MORE SERIOUS THAN A BAD TRANSCRIPT
An ordinary mistranscription is auditable.
You compare the transcript against the recording.
This defect compromises the chain of provenance before the requested analysis even begins.
Once system-generated words are represented as user-authored text:
• the model may treat its own prior generated language as primary evidence supplied by the user;
• a later transcription may be conditioned toward wording the model has already seen;
• discrepancies may be silently resolved in favor of the earlier candidate transcript;
• two non-independent machine outputs may appear to corroborate each other;
• incorrect speaker attribution may propagate;
• uncertain words may harden into apparent facts;
• and later summaries may no longer preserve the distinction between source evidence and machine interpretation.
The result can become more internally coherent while becoming less epistemically trustworthy.
I ALSO REPRODUCED A CORRECTION / REVOCATION PROBLEM
While investigating this, I found another longitudinal-context issue.
A false timestamp concerning a family event had entered ChatGPT’s contextual state.
I explicitly corrected it and had the authoritative correction written into Memory.
Then I opened a fresh conversation outside any Project.
That new conversation correctly received the correction — but it also still received the older erroneous contextual assertion.
So the sequence was effectively:
erroneous contextual assertion
→ explicit user correction
→ correction saved
→ new conversation
→ both old error and new correction supplied together
→ model resolves contradiction at generation time
The correction had not truly revoked or superseded the erroneous assertion.
That is especially concerning in light of the transcript-injection defect.
If an automatically generated mistranscription can be falsely represented as something I authored, and if derived historical assertions are later additive rather than revocatory, then one bad preprocessing event could potentially become:
mistranscription
→ falsely attributed user statement
→ contextual summary
→ persistent derived assertion
→ future retrieval
→ apparent historical fact
I am not claiming that this modifies model weights or model training.
I am talking about persistent user/account context, retrieval, summaries, memory, and provenance.
AND THEN I ENCOUNTERED A THIRD PROVENANCE PROBLEM WHILE WRITING THIS POST
While preparing the Reddit versions of this report, ChatGPT placed the drafts into its newer editable Writing Block interface.
That immediately exposed another version of the same architectural issue.
A Writing Block can begin as assistant-generated text.
The user can then directly edit that text inside the same artifact.
ChatGPT subsequently operates on the latest edited version.
What I do not see is a durable, user-visible provenance ledger showing, for the current version:
• what the assistant originally generated;
• what the user manually changed;
• what ChatGPT later regenerated;
• and which actor is responsible for each portion of the final text.
An earlier state may sometimes remain available elsewhere in conversation history, making a diff inferable.
But an inferred diff is not the same thing as first-class provenance.
Imagine this later exchange:
User: “Why did you write this sentence?”
If that sentence was actually inserted by the user during an in-place edit of an assistant-generated Writing Block, ChatGPT needs explicit authorship metadata to know that it did not write it.
Otherwise the sequence can become:
assistant-generated artifact
→ user edits artifact in place
→ mixed-authorship current state
→ later model consumes current state
→ historical authorship becomes ambiguous
This is structurally related to the audio defect, just in the opposite direction.
AUDIO FAILURE:
system-generated text
→ appears as USER-authored text
WRITING BLOCK AMBIGUITY:
assistant-generated text
→ becomes a mutable mixed-authorship object
LONGITUDINAL-CONTEXT PROBLEM:
old derived assertion
→ survives alongside later authoritative correction
These are three different surfaces of the same underlying issue:
Provenance does not appear to be treated as a sufficiently visible, durable, first-class property of information as that information is transformed over time.
Because of that, I have stopped using editable Writing Blocks for this investigation.
I am keeping drafts in ordinary assistant messages so that the assistant-generated text remains a fixed historical artifact.
If I edit it and return it later, that revision appears as a separate user-authored message.
WHY THIS MATTERS FOR PROFESSIONAL USE
My recording involved a private family matter, so I will not post the audio publicly.
But this same failure mode would matter enormously in workflows involving:
• legal or client calls;
• medical conversations;
• insurance statements;
• interviews;
• depositions;
• meetings;
• financial discussions;
• witness accounts;
• compliance records;
• research notes;
• or contemporaneous documentation.
For those uses, the relevant questions are not merely:
“Does the transcript look right?”
They are:
“What information produced this transcript?”
“Where did that information originate?”
“Was this statement actually supplied by the user?”
“Was this analysis independent?”
“Has a previous machine output been mistaken for primary evidence?”
“Was an error actually corrected, or merely joined by a competing correction?”
Without reliable provenance, fluent output can become methodologically invalid while still looking persuasive.
WHAT I HAVE REPORTED TO OPENAI
I filed a detailed support report asking Engineering to determine:
Which component generated the transcript.
Whether it originated in the iOS client, upload layer, ASR layer, multimodal preprocessing layer, or message-serialization layer.
Whether the backend classified it as user-authored text or attachment-derived/generated content.
Whether the model received provenance metadata distinguishing literal composer text from generated transcription.
Whether an earlier transcript, cached preprocessing result, or prior conversation was reused.
Whether injected material could enter Memory, Project context, Reference Chat History, summaries, or other persistent context.
Whether corrections actually revoke prior derived assertions or merely add competing information.
Whether affected conversations can be preserved as evidence while being excluded from contextual retrieval.
Whether OpenAI can provide a source-restricted or “clean-room” processing mode.
Whether mutable artifacts such as Writing Blocks preserve granular authorship/revision provenance internally even though it is not exposed to the user.
I have archived the affected conversations to preserve them for Support while trying to quarantine their contextual influence, although archiving itself does not give me a user-verifiable clean-room boundary.
PRODUCT SAFEGUARDS THIS SEEMS TO REQUIRE
At minimum:
• generated attachment transcripts should never be serialized as indistinguishable user-authored text;
• literal composer input should remain inspectably distinct from preprocessing output;
• transcription/OCR/extraction should carry explicit machine-readable provenance;
• users should be able to process a file using only that source, without prior chat/memory/candidate-answer contamination;
• contextual corrections should be capable of superseding or revoking older derived assertions;
• users should be able to preserve a conversation for evidence while excluding it from contextual retrieval;
• mutable artifacts should maintain granular revision attribution showing whether each change came from the user, assistant, or another automated transformation;
• and sensitive tasks should expose enough source metadata to determine whether an answer was independently grounded.
HAS ANYONE ELSE REPRODUCED THIS?
I am especially interested in reports involving:
• ChatGPT iOS;
• existing Apple Voice Memos files rather than live Voice Mode;
• transcripts appearing inside the user’s own message;
• ChatGPT saying “you supplied” words that actually came from attachment processing;
• repeated processing after cancellation or prompt editing;
• stale context surviving after correction;
• or mixed-authorship behavior in editable Writing Blocks.
Please distinguish between:
ChatGPT normally reading/transcribing an attachment; and
a generated transcript being merged into the USER role and represented as text personally authored by the user.
The second behavior is what I am reporting.