Clippy Vision

Clippy Vision

A fully local AI assistant that watches your screen to build context automatically, so you don't have to re-explain your work to an LLM every time. 100% private — nothing leaves your device.

🔗 Visit Clippy Vision
📁 AI & Machine Learning🗣️ English📅 September 3, 2026

Description

Anyone who works across a browser, an IDE, a terminal, and a document editor at the same time knows the friction of re-explaining what they're working on every time they open a new AI chat. Clippy Vision tries to remove that friction by having the AI simply watch what you're doing, instead of asking you to summarize it first — while keeping every bit of that observation on your own machine.

Clippy Vision is a fully local AI assistant that passively captures screen activity across applications, then reconstructs project context using a hierarchical memory system — raw events, session summaries, and long-term facts — with a fine-tuned classifier routing queries to the right memory tier and a ReAct agent using SQL tools to search it. It redacts its own window from captured screenshots, offers a system-tray toggle to pause capture entirely, and uses configurable retention (7-day raw events, 90-day summaries). Everything runs on-device with local LLM inference — there's no cloud component — and it's free and open source under the MIT license.

💬 Our review

The short version: Clippy Vision sits in the same space as Rewind.ai and Microsoft's Recall — continuous screen capture for AI recall — but its actual selling point is being fully local and open source where both of those are closed, cloud-adjacent products from companies with their own data incentives.

Rewind.ai popularized this category with a polished, commercially maintained product, but it still routes through Rewind's own infrastructure for some features and costs a subscription; Microsoft Recall ships built into Windows but has drawn real privacy criticism precisely because it's a first-party OS feature capturing everything you do. Clippy Vision's pitch is that neither trade-off is necessary — on-device inference, redaction of its own window, a hard tray-toggle kill switch, and MIT licensing mean you can audit exactly what it captures and where it goes (nowhere). The real cost is hardware: continuous local inference needs a reasonably capable machine, and as a younger open-source project it won't match Rewind's UI polish or search quality yet. Worth using if local-only processing is a hard requirement for you; Rewind.ai is still the smoother experience if you're fine with a closed, cloud-adjacent product from a funded company.

💰 Pricing

GratuitOpen source (MIT), aucun palier payant.

📊 Global score

58Average
🌐Availability15/100Faible

1 language · 0 platform

📄Profile100/100Excellent

Profile completeness

🤖 AI-enriched data

💰 Pricing model
🆓 Gratuit

Open source, licence MIT — aucun coût, s'exécute entièrement en local.

👥 Target audienceTravailleurs du savoir, développeurs et chercheurs qui veulent une IA avec mémoire contextuelle continue sans envoyer leurs données dans le cloud.
🗣️ Languagesen
🌍 Target countriesMarché anglophone, développeurs internationaux
👍

Pros

100% local, inférence LLM on-device, aucune donnée envoyée dans le cloud

Système de mémoire hiérarchique (événements → résumés → faits long terme)

Redaction automatique de sa propre fenêtre dans les captures

Bascule système tray pour désactiver la capture à tout moment

Gratuit et open source (MIT)

👎

Cons

Capture d'écran continue = surface de risque même si tout reste local

Nécessite une machine assez puissante pour l'inférence locale continue

Projet jeune comparé à Rewind.ai (produit commercial plus mature)

❓ Frequently asked questions

What is Clippy Vision in one sentence?
Is it free?
Does any of my data leave my device?
Can I turn off the screen capture?
Is it worth the money compared to alternatives?
Which tool should you pick for your case?