Open-source Python library that moves data from any source into your warehouse, handling schema and normalization for you.
Results for “data-engineering”
12 tools found
Real-time data movement platform: streaming, log-based CDC and batch through 200+ managed connectors, sub-100ms latency.
Open-source feature store that manages the machine-learning data teams use for both model training and real-time predictions.
Managed feature store and AI lakehouse platform for building production machine-learning systems with millisecond-latency feature serving.
Fully-managed change-data-capture platform that streams database changes into data warehouses in real time, without building Kafka/Debezium pipelines yourself.
Kafka-compatible streaming platform that runs diskless on cloud object storage (S3/GCS/Azure Blob), cutting Kafka's infrastructure cost sharply while keeping full API compatibility.
Platform for labeling, organizing and quality-checking the images, video and sensor data used to train computer vision and robotics AI models.
Data annotation platform built for organizations running many labeling projects at once across images, video, text, PDF and geospatial data.
Data annotation and evaluation platform that turns raw images, video, text and audio into labeled datasets for training and fine-tuning AI models.
Generates realistic fake versions of your production data — same shape and statistics, no real customer information — so developers can test against data that looks real without ever touching actual user records.
Developer-facing property data API covering nationwide US real estate records — ownership, valuations, comps, boundaries and skip tracing — for building proptech and fintech products.
Tool that watches every query hitting a company's data warehouse, catches bad data before it reaches a dashboard, and automatically writes the rule that prevents it from happening again.