Clippy Vision
A fully local AI assistant that watches your screen to build context automatically, so you don't have to re-explain your work to an LLM every time. 100% private — nothing leaves your device.
🔗 Visit Clippy VisionDescription
Anyone who works across a browser, an IDE, a terminal, and a document editor at the same time knows the friction of re-explaining what they're working on every time they open a new AI chat. Clippy Vision tries to remove that friction by having the AI simply watch what you're doing, instead of asking you to summarize it first — while keeping every bit of that observation on your own machine.
Clippy Vision is a fully local AI assistant that passively captures screen activity across applications, then reconstructs project context using a hierarchical memory system — raw events, session summaries, and long-term facts — with a fine-tuned classifier routing queries to the right memory tier and a ReAct agent using SQL tools to search it. It redacts its own window from captured screenshots, offers a system-tray toggle to pause capture entirely, and uses configurable retention (7-day raw events, 90-day summaries). Everything runs on-device with local LLM inference — there's no cloud component — and it's free and open source under the MIT license.
💬 Our review
The short version: Clippy Vision sits in the same space as Rewind.ai and Microsoft's Recall — continuous screen capture for AI recall — but its actual selling point is being fully local and open source where both of those are closed, cloud-adjacent products from companies with their own data incentives.
Rewind.ai popularized this category with a polished, commercially maintained product, but it still routes through Rewind's own infrastructure for some features and costs a subscription; Microsoft Recall ships built into Windows but has drawn real privacy criticism precisely because it's a first-party OS feature capturing everything you do. Clippy Vision's pitch is that neither trade-off is necessary — on-device inference, redaction of its own window, a hard tray-toggle kill switch, and MIT licensing mean you can audit exactly what it captures and where it goes (nowhere). The real cost is hardware: continuous local inference needs a reasonably capable machine, and as a younger open-source project it won't match Rewind's UI polish or search quality yet. Worth using if local-only processing is a hard requirement for you; Rewind.ai is still the smoother experience if you're fine with a closed, cloud-adjacent product from a funded company.
💰 Pricing
📊 Global score
🤖 AI-enriched data
Open source, licence MIT — aucun coût, s'exécute entièrement en local.
Pros
100% local, inférence LLM on-device, aucune donnée envoyée dans le cloud
Système de mémoire hiérarchique (événements → résumés → faits long terme)
Redaction automatique de sa propre fenêtre dans les captures
Bascule système tray pour désactiver la capture à tout moment
Gratuit et open source (MIT)
Cons
Capture d'écran continue = surface de risque même si tout reste local
Nécessite une machine assez puissante pour l'inférence locale continue
Projet jeune comparé à Rewind.ai (produit commercial plus mature)
