Managed feature store and AI lakehouse platform for building production machine-learning systems with millisecond-latency feature serving.
Results for “model-hosting”
44 tools found
Serverless cloud platform that runs Python code, including AI model training and inference, on GPUs with sub-second startup and pay-per-second billing.
Cloud inference platform for running and fine-tuning open-source language models fast, without owning any GPUs.
Enterprise cloud built specifically around NVIDIA GPUs, providing large-scale compute for training and running the biggest AI models.
GPU cloud built for AI research and production, scaling from a single rented GPU up to superclusters with over 165,000 GPUs.
Pay-by-the-second GPU cloud for training and running AI models, with no contracts and fast serverless cold starts.
A marketplace of ready-to-call AI models for generating images, video, audio and 3D content — like an app store of generative AI models you access through a single API instead of hosting each model yourself.
A European, publicly-traded cloud built specifically for AI workloads — GPU clusters, training, inference — for companies that want serious AI compute without depending on a US hyperscaler.
A safety checker for AI systems that catches hallucinations, unsafe answers and factual mistakes before they reach a user, using its own purpose-built judge models instead of relying on a general-purpose LLM to grade itself.
A single API endpoint that talks to 600+ AI models from 30+ providers, so switching from one AI model to another — or spreading requests across several for cost or reliability — doesn't mean rewriting your app's code.
A control panel for getting AI models and agents from a developer's laptop into real production use — deployment, scaling, and governance — built to run on whichever cloud a company already uses instead of locking them into one.
An AI inference API built specifically for speed — over 1,000 image, video, audio and language models served through one endpoint with sub-1-second latency and no cold starts.
Open-source on-device AI runtime that runs LLMs, vision models and speech models directly on phones, laptops and wearables — no cloud round-trip required.
Arkor lets your AI coding assistant write the training code for a custom AI model, then handles the expensive GPU work behind the scenes — so fine-tuning a model no longer requires being a machine-learning engineer.
BaseRT runs AI language models directly on your Mac's own chip instead of a data center — once it's running, there's no per-message bill because there's no server in the loop at all.
AI models on a phone-plan-style flat rate instead of a taxi meter — pay one predictable monthly price for a set number of requests across 40+ AI models, rather than watching a per-token bill climb with every word generated.
Lets an AI coding assistant (like Claude Code or Cursor) run on open-weight language models instead of a big-name API, with a guarantee that none of your code ever gets stored or used to train anything — like renting a private, no-questions-asked engine r
A service that trains a small, specialized AI model tailored to one specific task instead of using a giant general-purpose model for everything, at a fraction of the usual cost.
A free tool that lets a regular desktop computer run huge AI models that would normally need a room full of expensive server GPUs, by splitting the work smartly between your CPU and your (much smaller) GPU.
A service that figures out the fastest, cheapest way to run an open-source AI model in production for you — picking the right GPU, serving engine and settings automatically — instead of you spending weeks tuning it yourself.