Can Claude watch videos? Yes. Here's how, on your own machine
Last reviewed 1 October 2026
Not on its own. Paste a YouTube link into Claude and it cannot play the video or hear it. With a free skill called watch-for-me, Claude Code can: the video is transcribed on your own machine, the keyframes go to Claude as images, and you get a timestamped timeline of what was said and what was shown.
This page covers what you get, how to install it in Claude Code, Codex or any Agent Skills host, and what it costs in time and tokens. watch-for-me is made by Deepmark, the company behind this site, so the one section about Deepmark is clearly marked.
The short answer
Claude reads text and images. It does not take video or audio as input, so a video link on its own gives it very little: at best the page around the video, the title and description, and nothing that was actually said or shown.
Claude Code is different because it can run programs on your computer. watch-for-me is an agent skill, a folder of instructions plus a script, that gives it the missing senses:
- Hearing. The audio is transcribed on your device by open speech models, with timestamps. No captions needed, no cloud.
- Sight. One frame per slide, scene or code change is pulled out and tiled into labelled contact sheets that Claude reads as images.
- Several at once. Up to 10 links per call, YouTube, Instagram Reels, TikTok, X or local files, processed in parallel.
So the honest answer to "can Claude watch YouTube videos?" is: not in the chat app, yes in Claude Code with this skill installed.
What you get back
The default answer is a short summary, a timeline of what was said and shown, and any key text that appeared on screen. For YouTube, the timestamps come back as links that jump to that moment. Here is an abridged example of the shape, for a 1 hour talk:
Default mode, abridged example
**[1hr Talk] Intro to Large Language Models**, Andrej Karpathy, 59:48
Summary
- An LLM is two files: a parameters file and a short program that runs it
- Pretraining compresses a large slice of internet text into the parameters
- Fine-tuning on curated conversations turns that into an assistant
- The talk frames LLMs as the kernel of a new kind of operating system
- It ends on attacks: jailbreaks, prompt injection, data poisoning
Timeline
[00:20] LLM inference · slide: the two files of a 70B model
[04:17] Training · slide: GPUs, days and dollars for one training run
[14:14] Finetuning into an assistant · slide: a sample training conversation
[27:43] Tool use · demo: browser, calculator, Python, image generation
[42:15] LLM OS · diagram: the LLM as kernel, context window as RAM
[46:14] Jailbreaks · screenshots of roleplay and encoding tricks
…Ask for a different shape with a flag: a 3-sentence TL;DR, a plain words explanation, a step-by-step checklist, the on-screen code as files, verbatim quotes, or the answer to one question with timestamp evidence. Or skip the summary entirely and ask for work: "watch this tutorial and add the same feature to my app".
Install it in Claude Code, Codex or any Agent Skills host
watch-for-me ships as a standard agent skill, the same format as other Claude Code skills, so one repository covers every host. Pick yours:
Claude Code
Run these two commands inside Claude Code, one after the other:
/plugin marketplace add shafkathullah/watch-for-me/plugin install watch-for-me@watch-for-meThen use it as /watch-for-me. Autocomplete may show it as /watch-for-me:watch-for-me.
Any Agent Skills host
In your terminal
npx skills add shafkathullah/watch-for-me -g --skill watch-for-meCodex
In your terminal
codex plugin marketplace add shafkathullah/watch-for-meIn your terminal
codex plugin add watch-for-me@watch-for-meBy hand
Clone the GitHub repository and copy its skills/watch-for-me/ folder to ~/.claude/skills/watch-for-me/.
Requirements
You also need uv and ffmpeg 5.1 or newer. Everything else installs itself.
uv (macOS and Linux)
curl -LsSf https://astral.sh/uv/install.sh | shffmpeg on macOS (Linux: sudo apt install ffmpeg)
brew install ffmpegThe first run downloads the speech models once: about 2 GB for English, about 3.5 GB if you also watch non-English videos. To fetch them up front instead of on your first video, run this in Claude Code:
/watch-for-me --setupHow to use it
Give it one link, several links, or a file, with an optional mode:
/watch-for-me https://www.youtube.com/watch?v=zjkBMFhNj_g/watch-for-me https://youtu.be/aaa https://www.instagram.com/reel/bbb/ ~/Movies/demo.mp4/watch-for-me https://youtu.be/xyz --ask "which GPU does he recommend?"You can also just talk to your agent: "watch this and tell me the steps", "what does she say about pricing in these three videos?".
| Flag | What you get |
|---|---|
| (none) | Summary, timeline of what is said and shown, key on-screen text |
| --tldr | 1 to 3 sentences plus 3 key timestamps. Light visuals, fastest |
| --eli5 | Plain-words explanation, 150 words or fewer |
| --steps | A "You'll need" list and a numbered checklist with timestamps |
| --code | Transcribes on-screen code from full-resolution frames into files |
| --quotes | 5 to 15 verbatim quotes with timestamps, from the transcript only |
| --ask "question" | Answers the question with timestamp evidence |
| --from / --to | Only this part of the video; timestamps stay absolute |
| --hires | 1080p frames for small text and dense slides |
| --lang xx | Speech language (for example fr). Does not change the answer language |
| --cookies chrome | Uses your browser’s login for sites that need one |
| --playlist N | Accepts a playlist or multi-video post, takes the first N |
| --save | Also saves the link to your Deepmark library (paid, see below) |
Modes combine, for example --tldr --quotes. Up to 10 links per call.
How it works
- Download. yt-dlp fetches the audio and a 720p video stream in parallel, always the original-language audio, never an auto-dub. A 1 hour talk is about 50 MB.
- Transcribe on your device. English goes to NVIDIA’s Parakeet model, other languages to Whisper large-v3-turbo, chosen minute by minute by a small language detector. On Apple Silicon it runs on the GPU; other machines use a CPU build.
- Pick the keyframes. ffmpeg scene detection plus perceptual-hash deduplication keeps one frame per slide, scene or code change, then tiles them into labelled contact sheets. A 1 hour slide talk goes from 523 candidate frames to 73.
- Your agent reads. Subagents read the sheets in parallel while transcription is still running, zoom into full-resolution frames for small text and code, and the main agent merges everything into one timeline.
Everything is cached per video, so a second question about the same video starts in under a second.
Privacy: audio and transcription never leave your device. The frames your agent reads go to your agent’s model provider, like anything else you show it. No telemetry. The only network calls are the video download and the one-time model downloads from Hugging Face.
The limit: watch-for-me sees keyframes, not motion. It catches every slide, scene change and line of code on screen, but it won’t judge a golf swing.
Speed and token cost
Measured on the v0.1 build in September 2026, on an M1 Pro MacBook with 16 GB of memory and a roughly 5 MB/s connection, models already downloaded:
| Task | Measured |
|---|---|
| 1 hour slide talk: keyframes ready | 23 s after you ask |
| 1 hour slide talk: full transcript | 101 s (about 43x real time) |
| 3 videos at once: the two short ones | done in 3.5 s and 7.3 s |
| Same video again (cached) | under 0.2 s |
| Mixed English, French and Spanish audio | about 20x real time |
| CPU path (no Apple Silicon GPU) | about 19x real time on the same Mac |
| Claude Code, --tldr on a 2.5 minute video | 21 s end to end |
Tokens: visuals cost about 200 tokens per minute of a slide talk and about 3,000 per minute of a fast-cut video, and most of that is spent in subagents rather than your main conversation. A transcript is about 13,000 tokens per hour of speech. Long transcripts are digested by subagents, so a 1 hour talk costs the main conversation about 13,000 tokens (2 hours: about 18,000).
Reels, TikTok, X and local files: no captions needed
Reels, TikToks and X videos often have no captions you can copy, which is why "instagram reel transcript" and "tiktok transcript" are such common searches. watch-for-me never relies on captions: it transcribes the audio itself, so a Reel, a TikTok or a video posted on X gets the same transcript and timeline as a YouTube talk.
- X. Public video posts work without a login. A post with several videos asks you to pick:
--playlist 2takes the first two. - Instagram and TikTok. When a site asks for a login, add
--cookies chrome(or your browser) to use the session you already have. Chrome on macOS asks for Keychain access; Safari needs Full Disk Access for your terminal. - Local files. Pass a path instead of a link. A screen recording with no audio still gets a visual-only answer.
Save to Deepmark (optional)
This is the one part of the page about our product. Deepmark saves everything you bookmark (links, X posts, Reels, YouTube) and lets you search it in plain English, down to what was said in a video.
Add --save to any watch-for-me call and the video’s link and title also go to your Deepmark library. Only the URL and title are sent, never the transcript or frames. To enable it, connect Deepmark to Claude Code once, then run /mcp to sign in:
In your terminal
claude mcp add -s user --transport http deepmark https://usedeepmark.com/api/mcpIn claude.ai, add the same URL under Settings, then Connectors. Deepmark is a paid service, $7 a month billed yearly or $10 paid monthly, with no free tier: without a plan the save is refused with a message saying so. Everything else on this page works without it.
Common questions
Does it upload my video?
No. The video is downloaded to your machine and the audio is transcribed there; audio and transcripts never leave your device. The keyframe images your agent reads go to your agent’s model provider, like anything else you show it. There is no telemetry.
Which sites work?
YouTube, Instagram Reels, TikTok, X, Vimeo, Loom and most other sites yt-dlp supports, plus local video files. Sites that need a login (often Instagram) work with --cookies and your browser name. Playlists and multi-video posts need --playlist N. Live streams do not work; finished recordings of past streams do.
Does it need captions?
No. It transcribes the audio itself, so Reels, TikToks and X videos without captions work, and so do YouTube videos whose auto-captions are missing or wrong.
What do I need installed?
uv and ffmpeg. Python dependencies install themselves through uv, and the JavaScript runtime yt-dlp needs for YouTube ships with it. The fast path is a Mac with Apple Silicon on macOS 14 or later; Intel Macs and Linux use a slower CPU path, and Windows is untested. The first run downloads about 2 GB of speech models and runtime, about 3.5 GB if you also watch non-English videos.
Which agents support it?
Claude Code (tested), Codex and any host that supports Agent Skills and can run shell commands. Hosts without subagents read the contact sheets in the main conversation, and hosts without image input get transcript-only answers.
Is it free?
Yes. watch-for-me is MIT-licensed and runs on your machine, so there is nothing to pay for the skill itself. Your agent’s usage is billed by your agent’s provider as usual. The optional --save flag needs a paid Deepmark plan.
Can it see motion?
No. It sees keyframes, not motion. It catches every slide, scene change and line of code on screen, but it won’t judge a golf swing.
What does --save do?
It also saves the video’s link and title to your Deepmark library, where it becomes searchable by what was said and shown. Only the URL and title are sent, never the transcript or frames. It needs the Deepmark MCP connection and a Deepmark plan; without a plan the save is refused with a message saying so.
Give your agent any video
watch-for-me is free and MIT-licensed. Install it, paste a link, and your agent watches it on your machine.
Claude and Claude Code are trademarks of Anthropic, Codex of OpenAI, and YouTube of Google LLC; Deepmark is not affiliated with any of them. Numbers on this page are from the v0.1 measured pass on 29 September 2026. If something here is out of date, email support or open an issue on GitHub.