Description
When a company builds a chatbot or AI assistant, the scariest part usually isn't launching it — it's not knowing when it starts quietly failing. A tool call breaks, users have to repeat themselves, or the AI takes a bad path, and often nobody finds out until an actual customer complains. Selfship.ai was built by a team that hit exactly this problem running a 3-year-old AI trading chat product, and turned their fix into a product other teams can use.
Selfship.ai continuously watches every trace and multi-turn conversation of an AI agent in production, automatically detecting failure signals such as repeated tool-call errors, users having to reframe questions, or agents failing to reach the outcome the user wanted. It groups these failures by user intent, evaluates them, and closes the loop by generating and shipping a fix as a pull request — then checks afterward whether that fix actually worked.
💬 Our review
The short version: Selfship.ai's real pitch isn't detection (plenty of observability tools do that) — it's that it tries to close the loop by shipping the fix itself, then verifying the fix worked.
Generic LLM observability tools like LangSmith or Langfuse will show you traces and let you spot problems, and general error trackers like Sentry will catch crashes, but neither one groups failures by user intent or attempts an automated fix as a PR — that's Selfship's distinct claim. It's also a brand-new product (a single low-traffic Show HN launch at review time, no independently verified case studies, and no public pricing), so the automated-fix loop deserves real scrutiny before trusting it against a production agent: an early Show HN commenter asked what happens when Selfship proposes the wrong fix, and that risk is inherent to any system that writes code changes automatically. Worth a pilot if you're already drowning in silent agent failures and want to see the automated-PR loop in action; if you just need visibility without auto-remediation, a standard LLM observability tool is the safer, more proven choice for now.
💰 Pricing
📊 Global score
🤖 AI-enriched data
Aucune grille tarifaire publique au moment de la revue ; ouvert récemment en SaaS, tarifs à demander directement.
Pros
Détection automatique de signaux de défaillance fins (reformulations, échecs d'outils répétés)
Regroupement des échecs par intention utilisateur
Boucle fermée : propose et déploie un correctif en pull request
Vérifie après coup si le correctif a fonctionné
Cons
Produit très récent, aucune tarification publique
Peu de retours d'usage indépendants disponibles
Risque inhérent à la correction automatique de code (et si le correctif est faux ?)
Support on-prem non confirmé publiquement
