Open-source Python library that moves data from any source into your warehouse, handling schema and normalization for you.
Results for “etl”
113 tools found
Real-time data movement platform: streaming, log-based CDC and batch through 200+ managed connectors, sub-100ms latency.
Fully-managed change-data-capture platform that streams database changes into data warehouses in real time, without building Kafka/Debezium pipelines yourself.
Generates realistic fake versions of your production data — same shape and statistics, no real customer information — so developers can test against data that looks real without ever touching actual user records.
A data platform that pulls data from 200+ sources into a warehouse and lets you query it in plain English via an AI assistant.
Data-syncing platform that moves data both ways between a company's warehouse and the everyday tools (Salesforce, Google Sheets, Stripe) that teams actually work in.
A Python-native orchestration platform for running data pipelines, dbt workflows, and AI agents on either serverless or self-hosted infrastructure.
An all-in-one data platform that replaces a Fivetran + Snowflake + dbt + Looker stack with a single tool for ingesting, storing, and querying business data, plus an AI analyst that answers questions in plain English.
A single API that handles the whole messy world of PDFs for you — generating them, merging or splitting them, pulling structured data out of them, even translating them — so you don't have to stitch together five different libraries to do document work.
A streaming data platform, built by the people behind McLaren's Formula 1 data systems, that connects sensor data from R&D, testing, and production into one live feed instead of three disconnected systems.
Feed it a messy pile of PDFs, invoices or web pages and Datatera.ai turns them into clean, structured spreadsheet or CRM records — with an audit trail so a finance or legal team can trust where each number came from.
Decode GA4 automatically flattens Google Analytics 4's nested BigQuery export data into clean, ready-to-query tables, with usage-based pricing and no subscription.
An open-source SDK for building robotics data pipelines, managing the collection-to-delivery lifecycle of multimodal sensor data with quality checks and full provenance tracking, backed by Y Combinator.
A command-line tool that does to CSV, TSV and JSON what awk and sed do to plain text — filter, reshape and transform tabular data without spinning up a Python script.
Syncs Stripe, HubSpot, Postgres, and 40+ other business tools into one encrypted data lake you can query directly from Claude or Codex, without waiting on engineers or building a BI stack.
Open-source Python CLI that turns technical documentation into clean Q&A training datasets for LLM fine-tuning, using a local model.
Online spreadsheet cleaning tool that identifies and fixes messy data in Excel, CSV, and Google Sheets files. Users upload files for a free preview showing detected issues, then pay to unlock fully cleaned versions.
Open-source self-hosted tool for real-time data replication across multiple database engines (PostgreSQL, MySQL, MongoDB, Redis, SQLite, MariaDB). Eliminates need for Kafka or Debezium with lightweight synchronization via CDC, polling, or backfill.
Open-source tool that reads your existing dbt project, database schema, dashboards, and query logs, and automatically builds a structured map of your data's meaning — so an AI analytics agent has real context instead of guessing what a column means.
EU-hosted content extraction API that turns any webpage into clean markdown or plain text, built for developers feeding web content into RAG pipelines and AI applications.