Paste a video link and your Claude instantly knows everything in it.
DFYMD watches the whole video for you, every word, every slide, every screen, distills it into clean markdown, and writes what matters straight into your Claude's context, so a two-hour talk becomes something your AI just knows in minutes.
The problem and the fix
Stop taking notes on videos. Your Claude can do that.
The problem
You find a great talk, tutorial, or demo. Actually using it means watching the whole thing, taking notes, and pasting it into your AI. That is an hour you do not have.
Paste a video link
YouTube, Vimeo, or any direct URL. Drop it into the dashboard or send it from your Claude environment via the API.
DFYMD watches it
Captions-first transcription means no unnecessary spend on speech-to-text when captions exist. On the Vision tier, it also sees the screen: slides, dashboards, and code blocks via scene-change keyframes and OCR.
Your Claude reads, judges, and writes
The distilled markdown goes to your own Claude instance. It cross-references your CLAUDE.md and memory files, keeps what is relevant to your actual projects, and writes it straight into your context.
It watches the screen, not just the audio
Captures what was shown, not just said. Slides, code blocks, and dashboard numbers all end up in the markdown.
It knows your projects
Cross-references your CLAUDE.md and memory files so you get relevance, not a raw transcript. Things you already know get skipped.
Your key, your privacy
BYOK: the thinking runs on your own Anthropic key, locally. We never see your key or your private project context.
Audio $0.015/min, Vision $0.03/min. First hour free. Pay-as-you-go packs that never expire, or a subscription that rolls over.
Live demo
One honest run, start to finish.
A simulated walk-through of the real journey: copy a link, ask your own Claude Code, watch the skill do the work, and see your project files get smarter. No signup needed to watch.
Processing tiers
Choose how deeply you want Claude to see.
Both tiers produce clean markdown. Vision goes further: it also captures what is shown on screen, not just what is said.
Audio
Full speech-to-markdown. Captures everything spoken: insights, frameworks, instructions, and key quotes. Ideal for podcasts, interviews, and talk-based content.
- High-accuracy transcription
- Structured markdown output (headings, bullets, callouts)
- Cross-referenced key concepts
- Speaker attribution when detectable
Vision
Everything in Audio, plus frame-by-frame visual analysis. Reads slides, code on screen, diagrams, and whiteboard content. Best for tutorials, demos, and technical walkthroughs.
- Everything in Audio
- On-screen slide and code extraction
- Diagram and chart description
- Visual and speech cross-referencing
- Fidelity-matched timestamps
Pricing
Pay as you go, or subscribe and save.
Packs never expire. Subscriptions refresh your wallet monthly. All plans start with free processing time.