A free, open-source Rust library that quickly figures out whether a PDF is text-based or scanned and pulls out clean, structured text — so you only pay for expensive OCR on the PDFs that actually need it.
#data-engineering
47 tools curated in this category — including pdf-inspector, HFlow, Hubble
techFind on mySelectas all sites and tools related to data-engineering. This selection of 47 resources is reviewed and maintained by the community. The most popular include pdf-inspector, HFlow, Hubble. Each tool comes with a review, tags, comparisons and alternatives to help you make the best choice.
An open-source SDK for building robotics data pipelines, managing the collection-to-delivery lifecycle of multimodal sensor data with quality checks and full provenance tracking, backed by Y Combinator.
An API that assembles a patient's full medical record across every provider, using EHR connections plus browser and voice agents as fallback.
Type in a search term like "plumbers in Chicago" and get back a spreadsheet of every matching business on Google Maps — names, phone numbers, emails, ratings — instead of manually copying details out of Google Maps one listing at a time.
A Python-native orchestration platform for running data pipelines, dbt workflows, and AI agents on either serverless or self-hosted infrastructure.
Data-syncing platform that moves data both ways between a company's warehouse and the everyday tools (Salesforce, Google Sheets, Stripe) that teams actually work in.
A tool that automatically fills in the missing blanks in your sales database — company size, job title, whether they just raised funding or started hiring — so your sales team spends time talking to prospects instead of googling them one by one.
A data platform that lets an AI agent or application query all your different databases and data sources with one SQL question — combining search, filtering, and even AI model calls right inside the query — without the usual months-long project of buildin
Tool that watches every query hitting a company's data warehouse, catches bad data before it reaches a dashboard, and automatically writes the rule that prevents it from happening again.
Developer-facing property data API covering nationwide US real estate records — ownership, valuations, comps, boundaries and skip tracing — for building proptech and fintech products.
Generates realistic fake versions of your production data — same shape and statistics, no real customer information — so developers can test against data that looks real without ever touching actual user records.
Data annotation and evaluation platform that turns raw images, video, text and audio into labeled datasets for training and fine-tuning AI models.
Data annotation platform built for organizations running many labeling projects at once across images, video, text, PDF and geospatial data.
Platform for labeling, organizing and quality-checking the images, video and sensor data used to train computer vision and robotics AI models.
Kafka-compatible streaming platform that runs diskless on cloud object storage (S3/GCS/Azure Blob), cutting Kafka's infrastructure cost sharply while keeping full API compatibility.
Fully-managed change-data-capture platform that streams database changes into data warehouses in real time, without building Kafka/Debezium pipelines yourself.
Managed feature store and AI lakehouse platform for building production machine-learning systems with millisecond-latency feature serving.
Open-source feature store that manages the machine-learning data teams use for both model training and real-time predictions.
Real-time data movement platform: streaming, log-based CDC and batch through 200+ managed connectors, sub-100ms latency.
Open-source Python library that moves data from any source into your warehouse, handling schema and normalization for you.