ChatGPT Thinking Animation: Searching and Reasoning
The most interesting part of an AI conversation video is often not the answer. It is the pause before it.
A prompt lands, and for two or three seconds the interface tells you something is happening — a pulsing dot, a spinning loader labelled "Searching the web", a line that says "Reasoning". That gap is where the viewer leans in. It is also the part most fake ChatGPT videos skip entirely, cutting straight from prompt to fully-formed answer, which is exactly why those clips feel wrong without the viewer being able to say why.
This guide covers the states that happen before and around the response in the ChatGPT template: the thinking pulse, the labelled activity indicators, error states, and the timing values that make a sequence read as real. If you want the text-streaming side specifically, the streaming text effect guide covers that in depth.
First, what this actually is
Worth stating plainly, because it shapes how you should use the format: MockClip does not run a language model. It is a deterministic renderer. You write the prompt, you write the response, and the template animates the interface around them — including the activity states that a real assistant would show while working.
That is a feature for video work rather than a limitation. A scripted exchange renders identically every time, so you can iterate on pacing without a model producing a different answer on each take, and you can demonstrate a specific workflow without waiting on a live response. It does mean the honesty burden sits with you: if you are demoing a product, the response you write should reflect what the product genuinely does.
The two waiting states
There are two distinct ways the template fills the gap before a response, and choosing between them is the main creative decision in this format.
The thinking pulse. A white circle that scales gently in a loop. It carries no text. It appears during the delay before an assistant message when that message has no activity attached. This is the neutral wait — the viewer knows something is coming but not what kind of work is happening.
The activity indicator. A spinning loader paired with a text label. It replaces the pulse and tells the viewer specifically what is supposedly happening: searching, reasoning, reading a file. This is the narrated wait.
The rule of thumb: use the pulse when the wait is just pacing, and use an activity indicator when the wait is content. If the point of your video is that the assistant went and looked something up, the label is the point. If the assistant is answering from general knowledge, a labelled indicator would be a lie about what happened, and the pulse is the honest choice.
The activity types
Seven activity types are available, and each has an implied meaning that viewers read instantly:
searching— going out to the web for current information. The most legible to a general audience.browsing— reading a specific page rather than running a query. Subtly different from searching, and useful when your script names a source.reasoning— working through a problem rather than retrieving. Suits multi-step or analytical prompts.analyzing_image— processing an image the user supplied.generating_image— producing an image. Pairs with the template's inline image support, covered below.reading_file— working through an uploaded document.custom— your own label text, for anything the list above does not cover.
Each activity takes an optional label and a duration. The label overrides the default text, which matters more than it sounds: "Searching the web" and "Searching 4 sources" imply different things about thoroughness, and both are a single field.
The custom type is the one worth remembering. If your video is about a specific tool or workflow — checking a calendar, querying a database, running a test suite — a custom label represents it directly instead of forcing it into the nearest built-in category.
Browser-based. No install.
Open the ChatGPT templateTiming the wait
Three values control pacing, and all three are set per message rather than globally.
delayBefore is how long to wait before the message begins. For an assistant message with no activity, this is how long the thinking pulse runs. Configurable up to thirty seconds.
action.duration is how long the activity indicator holds. Two to four seconds is the realistic range for searching or reasoning. Shorter than a second and the viewer cannot finish reading the label, which wastes the beat entirely — a common and invisible mistake, because the creator already knows what the label says.
streamingSpeed controls how fast the response types out, and is also per message. A one-line answer can appear briskly while a longer explanation streams slowly enough to read. Leaving every message at the same speed is the equivalent of reading a script in monotone.
A pattern that works for a single-exchange clip: a short delayBefore of under a second after the prompt, an activity indicator for three seconds, then a moderate streaming speed for the response. The total is around six to eight seconds before the answer completes — long enough to feel like work happened, short enough for a vertical feed.
For longer content, vary it. If your video has three exchanges, do not give all three the same three-second search. Make one instant, one slow, one labelled differently. Uniform pacing across exchanges is what makes a multi-turn clip feel synthetic.
Reasoning as its own beat
The reasoning indicator deserves separate treatment because it changes what a viewer expects from the answer that follows.
A searching label promises retrieved facts. A reasoning label promises a conclusion — and a viewer who has watched a reasoning indicator for three seconds expects the response to contain judgement, not a list. If you label the wait as reasoning and then deliver a bare definition, the mismatch registers as disappointing even though the viewer will not articulate why.
This pairing — indicator promises, response delivers — is the most useful discipline in the format. Match the label to the shape of the answer.
Error states
Any message can carry an error string, which renders as a red error box with a warning icon in place of a normal response.
This is a narrow feature with a few strong uses. Content about rate limits, failed requests, or context-length problems needs to show the failure, and a screenshot cannot show the timing — the wait, then the failure. Troubleshooting tutorials benefit for the same reason. It is also the honest way to depict a limitation: if your video is about what an assistant cannot do, an error state says so directly.
A sequence that works well: prompt, activity indicator for a few seconds, then the error. The wait makes the failure land harder, because the viewer has already invested attention in an answer that never arrives.
Images and the generation states
The template supports inline images on a message, with a position of before or after the text and a configurable reveal duration. Paired with the generating_image activity, this produces the full arc: the indicator runs, then the image appears with a blur-to-sharp reveal, then the text streams — or the reverse, depending on the position you choose.
Position matters for pacing. Setting the image before the text means the visual lands first and the text explains it. Setting it after means the text builds expectation and the image pays it off. For short-form video, after usually holds attention better, because the wait for the image is doing work.
The ChatGPT image generation animation guide covers the reveal effect in more detail.
The chrome around the conversation
The activity states sit inside an interface, and a handful of clip-level settings change what that interface implies before a single message appears.
The model label. The template takes a model name as free text, shown in the header. This is a small field with an outsized effect on framing — a clip labelled with a reasoning-tier model sets a different expectation for a three-second reasoning indicator than one labelled with a fast, lightweight model. Keep it consistent with the story you are telling.
Incognito mode. Toggling incognito switches the interface to the temporary-conversation state, including the privacy notice on the welcome screen and a changed input placeholder. Useful when the framing of your video is about privacy, sensitive prompts, or throwaway sessions.
The welcome screen. You can show or hide the welcome state, set your own welcome text, and toggle the quick-action buttons beneath it. There is also an option to keep the welcome visible rather than clearing it when the conversation starts. For a cold open, starting on the welcome screen and having the prompt land gives you a half-second of recognisable context before anything happens — the viewer knows exactly what they are looking at.
Theme. Light and dark, as with every MockClip template. Dark is the more common choice for short-form because it holds contrast against captions, but light reads as more ordinary and workaday, which suits tutorial content.
None of these are load-bearing on their own. Together they set the frame that the activity indicators play out inside, and leaving all of them at default is a missed half-second of storytelling.
Where this format fits
- Product and workflow demos. Script the exact exchange your product enables. Because the render is deterministic, you can iterate on wording and pacing without a live model producing different output each take. See ChatGPT conversation video for marketing demos.
- Prompting tutorials. Show a prompt and the shape of a good response. The activity indicator communicates that the assistant did work, which is often the teaching point.
- Explainer content about how AI tools work. The activity states are the most legible way to show that an assistant retrieves, reads, and reasons rather than simply knowing.
- Troubleshooting and limitation content. Error states, covered above.
- Cold opens. A prompt and a three-second search indicator is a complete hook in under five seconds, with the answer withheld until after your intro.
Common mistakes
Skipping the wait entirely. Cutting from prompt to complete answer is the most common error in fake AI conversation videos. It removes the beat that makes the exchange feel like an exchange.
Indicators too short to read. Under a second, the label is decorative. You know what it says; the viewer does not.
Labels that do not match the answer. A reasoning indicator followed by a one-line retrieved fact, or a searching indicator followed by an opinion. The mismatch registers even when the viewer cannot name it.
Uniform pacing across every exchange. Identical delays and streaming speeds across a multi-turn clip flatten it. Vary at least one value per exchange.
Overusing activity indicators. If every message is labelled, none of them is a beat. Use the plain thinking pulse for the neutral waits so the labelled ones stand out.
Writing responses that oversell. Since you author both sides, it is easy to write an assistant response that claims more than your product does. Keep it accurate.
Quick start
- Open the ChatGPT template
- Write the user prompt as the first message
- Add the assistant response and attach an activity to it
- Pick the type —
searching,reasoning, orcustomwith your own label - Set the activity duration to around three seconds
- Set a streaming speed that lets the response be read at natural pace
- For a failure sequence, add an error string instead of a response
- Preview, then export the MP4
Export tiers and watermark options are listed on the pricing page.
Related templates and guides
- ChatGPT template — the editor used in this guide
- How to create a fake ChatGPT conversation video — the format pillar
- ChatGPT streaming text animation effect — the text-streaming layer
- ChatGPT image generation animation — the blur-to-reveal image effect
- ChatGPT conversation video for marketing demos — the product-demo application
- Connect MockClip to Claude and ChatGPT — generating these clips through MCP
- All MockClip templates — the full format index
- MockClip vs CapCut — template approach versus manual editing
- Pricing — export tiers
Frequently Asked Questions
What is the ChatGPT thinking animation?
It is the state shown between a user sending a prompt and the response beginning to stream. In the MockClip template this renders as a pulsing circle that scales gently in a loop. It appears during the delay before an assistant message when that message has no activity indicator attached to it.
What is the difference between the thinking pulse and an activity indicator?
The pulsing circle is the generic waiting state and carries no label. An activity indicator replaces it with a spinning loader and a text label such as searching or reasoning. Use the pulse when the wait is neutral, and an activity indicator when you want the viewer to know what is supposedly happening.
Which activity states can I show?
The template supports searching, reasoning, analyzing image, generating image, reading file, browsing, and a custom type. The custom type lets you supply your own label text, which is how you represent an activity the built-in list does not cover.
Does MockClip actually search the web or run a model?
No. MockClip is a deterministic renderer. It animates the interface states of an AI conversation using the text you supply. Nothing is generated, no search is performed, and no model is called. You write both sides of the exchange and the template renders them.
How long should an activity indicator run?
Two to four seconds reads as realistic for searching or reasoning. Under a second and the viewer cannot read the label, so the beat is wasted. Beyond about five seconds, a short-form audience will scroll unless something else is on screen. The duration is configurable up to thirty seconds.
Can I show an error response?
Yes. Any message can carry an error string, which renders as a red error box with a warning icon instead of a normal response. This is useful for content about rate limits, failures, and troubleshooting.
Can I control how fast the response types out?
Yes. Each assistant message has its own streaming speed, so a short answer can appear quickly while a longer one streams at a readable pace. The delay before each message is also set per message.
Is this suitable for product demo videos?
Yes, and it is one of the more common uses. Because you author both the prompt and the response, you can script an exact demo without waiting on a live model or re-recording when the output varies. Keep the content honest about what your product actually does.
Related Articles
How to Create a Fake ChatGPT Conversation Video
The complete 2026 guide to creating realistic fake ChatGPT conversation videos for TikTok, YouTube Shorts, and Reels — no coding, no editing, free.
How to Recreate the ChatGPT Streaming Text Effect in Videos
Add the iconic ChatGPT word-by-word streaming text animation to your videos. No coding or screen recording needed.
Animate the ChatGPT Image Generation Effect (Blur Pulse to Reveal)
Recreate the ChatGPT AI image generation animation — blur pulse while generating, then sharp reveal — in video mockups with no screen recording.
ChatGPT Conversation Video for Product Marketing Demos
Create fake ChatGPT conversation videos for product demos and marketing. MockClip renders AI chat animations with streaming text as MP4 video.
How to Connect MockClip to Claude and ChatGPT (MCP Setup Guide)
Step-by-step guide to configure MockClip inside Claude Desktop and ChatGPT using the Model Context Protocol. Generate conversation mockup videos from plain-English prompts.