Data Engineering AI Startup Jobs
Explore startups tagged with Data Engineering and compare hiring activity, company profiles, and direct job links. This page is indexable only when a tag reaches at least 5 companies to avoid thin content.

Databricks
761 jobsUnified analytics platform for data and AI, helping companies process and analyze big data in the cloud.

Scale AI
191 jobsData infrastructure company providing high-quality training data for AI applications, recently partnered with Meta.

Hone Health
177 jobsHone Health is a direct-to-consumer telehealth clinic focused on longevity, hormone optimization, and preventive care for men and women. Clinicians run virtual visits, order and review lab biomarkers, and build personalized wellness plans. The company raised a $33M Series A in early 2025 and acquired women's-health telehealth startup ivee.

Grafana Labs
108 jobsCompany behind the open-source Grafana observability stack providing monitoring, logging, and tracing solutions, reaching $400M ARR as a fully remote company across 40+ countries.

Hightouch
68 jobsData activation platform (Reverse ETL, Customer Studio) to sync warehouse data to business tools.

Perle
62 jobsPerle is an AI training data platform that combines human expertise with adaptive workflows to help companies collect, annotate, and evaluate specialized training data for generative AI, LLMs, and RLHF. Their vetted global network of domain experts provides modular solutions for data annotation, enrichment, and adversarial robustness assessment.

Together AI
62 jobsTogether AI provides an AI-native cloud platform for inference, fine-tuning, GPU clusters, and large-scale model training, helping AI engineers and researchers build and scale open-source and enterprise AI systems.

Cribl
58 jobsData pipeline platform that gives you control over your observability data.

Encord
45 jobsEncord is a data development platform for AI that helps teams curate, manage, and label training data at scale. The company serves 300+ AI teams including Toyota and Skydio, processing over 5PB of data.

Snorkel AI
45 jobsSnorkel AI provides a data-centric AI platform for building specialized AI systems, helping enterprises and frontier teams create, curate, evaluate, and govern expert data for production models and agents.

Weka
44 jobsWEKA builds a cloud and AI data platform that accelerates model training and inference workloads with high-performance, software-defined storage.

Mecka
36 jobsMecka (Mecka AI) is a New York–based startup building the data engine for Physical AI — it collects real human motion data from body sensors and iPhones to train robots and embodied AI. The company provides the data, evaluation, and deployment infrastructure that helps robots learn from real human activity and move from lab demos to reliable real-world work.

Gruve
30 jobsGruve delivers AI-native infrastructure, inference systems, and enterprise AI agents for inference-heavy workloads with an emphasis on speed, security, and measurable outcomes.

AfterQuery
28 jobsAfterQuery is a San Francisco-based applied research lab that builds expert-curated AI training datasets and model evaluation suites for frontier AI labs. It leverages a network of nearly 100,000 domain experts (developers, attorneys, and other professionals) to generate high-quality data, evals, and dev environments. The company reached over $100M ARR within roughly 14 months of founding.

Distyl AI
28 jobsData intelligence platform that unifies messy operational data, applies AI agents, and routes insights back into business workflows.

Reducto
28 jobsReducto provides a high-quality AI document ingestion and parsing API for large language models. The Y Combinator-backed company processes nearly a billion pages monthly for leading AI teams like Harvey and Scale AI.

Hex
27 jobsHex is a collaborative analytics workspace that combines notebooks, SQL, data apps, and AI-assisted workflows for data teams.

Anaconda
25 jobsAnaconda provides the world's most popular open-source Python and R distribution for data science and AI development. Serving over 45 million users, its platform enables enterprises to manage packages, environments, and AI workflows at scale with security and governance controls.

Fundamental
25 jobsFundamental builds large tabular models and enterprise AI infrastructure for prediction and analysis on complex business data, focused on tabular reasoning and decision support.

Astronomer
23 jobsCompany behind Astro, a managed Apache Airflow DataOps platform for data & AI pipelines.

Encharge AI
23 jobsDeveloping analog in-memory compute chips and software for energy-efficient AI at the edge.

Omni
21 jobsOmni is a modern business intelligence and analytics platform that combines a unified semantic data model with SQL flexibility, enabling AI-powered trustworthy answers in seconds. The platform supports embedded analytics, custom dashboards, and governed data exploration.

Nexthop AI
20 jobsNexthop AI builds networking systems for AI-scale data centers, focusing on high-performance switching infrastructure for hyperscale and cloud environments.

Anyscale
19 jobsAnyscale provides a fully-managed compute platform built on Ray, an open-source distributed computing framework originally developed at UC Berkeley. The company enables developers to build, deploy, and scale AI and Python applications without requiring deep infrastructure expertise.

Eon
18 jobsEon is the first cloud backup posture management (CBPM) platform, automating and unifying complex cloud backups into a queryable data lake for fast recovery, compliance, and AI analytics. Founded by the team behind AWS Disaster Recovery, Eon converts idle backup data into an accessible secondary storage layer for enterprise AI workloads.

Prior Labs
17 jobsPrior Labs builds tabular foundation models that understand spreadsheets and databases, enabling instant pattern inference across any dataset without task-specific training. Their flagship model TabPFN, trained on 130 million synthetic datasets, ranks #1 on the TabArena benchmark and scales to 10 million rows, serving Fortune 500 companies like Hitachi.

Domino Data Lab
15 jobsDomino Data Lab provides an enterprise AI and MLOps platform for model development, deployment, governance, and compliance. Trusted by over 20% of the Fortune 100, the platform helps data science and machine learning teams manage regulated model workflows at scale. The company has rebranded to Domino and is backed by Sequoia Capital, Coatue, Great Hill Partners, and Highland Capital Partners.

Firecrawl
14 jobsFirecrawl is a web data infrastructure platform that converts websites into clean, structured data optimized for AI applications through a simple API, turning entire websites into LLM-ready markdown or structured data.

Databento
13 jobsDatabento is a market data platform that delivers real-time and historical data on futures, options, and equities through a single API, positioning itself as a developer-first, usage-based alternative to legacy providers like Bloomberg and LSEG. Founded in 2019 by former quant trader Christina Qi, it lets trading firms, funds, and fintechs access normalized institutional-grade market data without long-term contracts.

LlamaIndex
13 jobsLlamaIndex is a data framework for LLM applications that enables developers to connect, index, and query custom data sources with large language models through their open-source library and LlamaCloud platform.

PhaseV
12 jobsML-driven adaptive trials and clinical development optimization.

Protege
12 jobsProtege operates a governed marketplace platform for ethical sourcing of multimodal, real-world AI training data with compliant data exchange capabilities.

San Francisco Compute
12 jobsSF Compute provides rentable, large, low-cost GPU clusters for AI pre-training workloads. The platform operates as a marketplace connecting AI teams with on-demand high-performance computing capacity, offering flexible access to supercomputing-scale infrastructure with InfiniBand interconnects.

TinyFish
12 jobsTinyFish provides enterprise web agents that automate complex web-based workflows and extract structured data from websites at scale. The platform enables Fortune 500 companies like Google and DoorDash to automate web interactions, streamline data collection, and integrate web automation into their business processes.

Bespoke Labs
11 jobsBespoke Labs is an applied AI company that builds the reinforcement learning environments, data curation systems, and infrastructure used to train, evaluate, and post-train reliable long-horizon AI agents for frontier labs and enterprises. It is known for open-source tooling like Bespoke Curator and open datasets. Founded in 2024 by Mahesh Sathiamoorthy and Alex Dimakis, it is headquartered in Mountain View, California.

Stedi
11 jobsStedi is the only programmable, API-first healthcare clearinghouse, replacing legacy batch-and-portal clearinghouses with a modern cloud-native platform. It processes over 1 billion healthcare transactions per year.

Junction
10 jobsJunction (formerly Vital) modernizes healthcare infrastructure with seamless lab testing and device data integration, connecting over 500 wearables and medical devices with 10+ lab networks including Labcorp and Quest across all 50 states.

Rowspace
10 jobsRowspace is a San Francisco-based AI platform that helps financial services firms—especially private equity and investment firms—turn proprietary, scattered internal data into actionable decision intelligence. It unifies information across document repositories, investment platforms, and accounting systems to support investment research, risk analysis, and deal evaluation. Founded in 2024 by former Notion CTO Michael Manapat and CFO Yibo Ling, it emerged from stealth in February 2026.

Datology AI
9 jobsAI training data curation platform helping enterprises optimize ML training data at petabyte scale.

David AI
9 jobsDavid AI is the world's first dedicated audio data research lab, building the data layer for next-generation audio AI. Founded by former Scale AI engineers, serving most FAANG companies and major AI labs.

MotherDuck
9 jobsMotherDuck is a serverless cloud data warehouse built on the open-source DuckDB engine, enabling fast SQL analytics with no infrastructure to manage. The platform supports hybrid local-cloud execution, allowing analysts to query data seamlessly across laptop and cloud.

Pulse
9 jobsPulse is a production-grade unstructured document extraction platform that converts complex PDFs, Word, Excel, and other document formats into structured, LLM-ready data. Built for enterprise reliability with OCR, bounding boxes, and vision language model capabilities.

Sapien
9 jobsSapien builds AI-native analysts for finance and operations teams, connecting ERP, warehouse, spreadsheet, and operational data. Its agents help CFO and analytics teams find profit drivers, explain variance, and act on messy transaction-level data faster.

Unstructured
9 jobsOpen-source data preprocessing platform that extracts, cleans, and transforms unstructured documents (PDFs, images, HTML, emails) into structured formats optimized for AI and LLM pipelines.

Novig
8 jobsNovig operates a peer-to-peer sports prediction market where users trade directly with one another instead of against the house, eliminating hidden fees and biased odds. Founded in 2021 by Jacob Fortinsky and Kelechi Ukah and backed by Y Combinator (S22), it has become the fastest-growing platform in the U.S. sports prediction category. In February 2026 it raised a $75M Series B at a $500M valuation to compete with Kalshi and Polymarket.

Novellia
6 jobsNovellia is a patient-powered health data platform that lets anyone in the US access and unify up to ~20 years of their complete medical history in under 30 seconds, for free. It turns that de-identified, patient-consented data into real-world data products for biopharma R&D, already working with several of the top 10 pharmaceutical companies. The company offers both a web platform and a patient-facing mobile app.

Relace
6 jobsRelace is a provider of auxiliary coding models for faster, more reliable AI code generation that makes it easy to deploy production-ready coding agents with models co-optimized with infrastructure to achieve state-of-the-art performance across million-line repositories.

Unlimited Industries
6 jobsUnlimited Industries is an AI-native construction company that vertically integrates design and build for large-scale infrastructure projects including data centers, energy facilities, and advanced manufacturing. The company's proprietary AI platform can explore tens of thousands of design configurations to optimize costs and timelines, reducing pre-construction engineering from months to weeks. Founded by serial entrepreneurs and backed by Andreessen Horowitz, Unlimited is rethinking how America's critical infrastructure gets built.

Covariance
5 jobsCovariance is a machine-learning platform that turns external and alternative data into firm-level KPIs and business forecasts. Its product analyzes actual purchase data from millions of households to give consumer brands insights on category shifts, competitor moves, audience behavior, and campaign impact, and also serves investment users with alternative-data signals. Founded out of MIT research by CEO Michael Fleder, it is based in New York City.

Graphon AI
5 jobsGraphon AI is a pre-model intelligence layer that transforms enterprise multimodal data (documents, video, audio, logs, databases) into persistent relational memory using graphon mathematical functions. It enables LLMs and AI agents to reason accurately across unlimited connected data sources with larger effective context. The company was founded by researchers from Amazon, Meta, MIT, Google, Apple, and NVIDIA.

Prefect
5 jobsPrefect builds workflow orchestration and AI infrastructure software that helps teams automate, observe, and manage data and application workflows.

Tracer
5 jobsTracer is the first pipeline monitoring system purpose-built for high-performance computing in life sciences, providing real-time performance metrics, cost breakdowns, and optimization insights for complex computational pipelines.

Colossal Biosciences
4 jobsColossal Biosciences is a genetic engineering and de-extinction company using CRISPR technology to restore extinct species like the woolly mammoth and protect critically endangered ecosystems.

Parabola
4 jobsParabola is an AI workflow builder for operations and finance teams that turns messy operational data from spreadsheets, PDFs, emails, APIs, and databases into repeatable automated workflows. Teams use it to automate reconciliations, inventory and order operations, reporting, and other manual business processes without relying on engineering support.

Spiral
4 jobsSpiral is a data infrastructure company that provides a multimodal data platform for AI, unifying governance and exposing a single API for every data modality including video, audio, geospatial, and text, engineered for machine-scale throughput to keep GPUs fully saturated.

Transcend
4 jobsTranscend is an enterprise-grade data privacy infrastructure platform that serves as the compliance layer for customer data. It enables organizations to automate data subject requests, map data across systems, manage consent, and activate data for AI responsibly at scale.

ZeroEntropy
4 jobsZeroEntropy provides a high-accuracy search API over unstructured data for AI agents and RAG applications. The YC-backed company builds smarter retrieval models enabling AI agents across healthcare, law, and sales.

Artie
3 jobsFully managed change data capture (CDC) streaming platform that replicates production databases into data warehouses and lakes in real time. Trusted by Substack, ClickUp, and Alloy, processing over 700 billion rows annually.

Deepnote
3 jobsCollaborative cloud data notebook platform for data science and analytics teams.

Tensormesh
3 jobsSemantic KV caching layer built for LLM inference, enabling AI applications to reduce inference costs and latency by reusing cached computation across similar prompts.

Alloy
2 jobsAlloy is a data platform for robotics that helps companies process, organize, and search through the massive volumes of sensor, camera, and telemetry data their robots generate. The Sydney-based startup enables natural language search across robot data and automated issue detection, reducing data processing time by up to 90%.

Blockworks
2 jobsBlockworks provides crypto market data, research, and media for digital asset markets. It operates an onchain capital markets intelligence and networking platform that connects investors and businesses, and delivers breaking news and premium insights about digital assets and web3.

Credal.ai
2 jobsCredal provides a secure AI agent platform for enterprises, enabling teams to build AI agents and MCP-connected workflows across internal data sources with governance controls.

Radical AI
2 jobsRadical AI is building next-generation AI infrastructure with a focus on custom silicon and compute systems optimized for large-scale model training and inference. It aims to reduce the cost and energy consumption of AI workloads through purpose-built hardware and software co-design.

Shovels
2 jobsShovels builds construction intelligence software that turns fragmented building permit data into actionable market and go-to-market signals through APIs and analytics tools.

Syenta
2 jobsSyenta develops Localized Electrochemical Manufacturing (LEM) technology for advanced semiconductor chip packaging, enabling scalable, high-density interconnects without traditional lithography. Spun out from the Australian National University, their approach addresses memory bandwidth bottlenecks in AI computing.

OneSchema
1 jobsAI-driven CSV and PDF data import automation platform for seamless customer onboarding.

Pytho AI
1 jobsProvides a unified interface to design AI workflows by connecting data, models, and automations.

Structify
1 jobsAI-powered data platform that transforms unstructured web data and documents (websites, PDFs, pitch decks, reports) into structured, enterprise-ready datasets using their proprietary DoRa model that navigates and extracts data like a human, enabling real-time web extraction for business intelligence and data workflows.

Tinybird
1 jobsTinybird is a real-time data platform that enables data and engineering teams to build real-time data products and APIs at scale. The platform ingests, transforms, and serves large volumes of data with sub-second latency for analytics and operational intelligence.

Tonic AI
1 jobsGenerates realistic synthetic data to power software testing and analytics without exposing sensitive production data.

Anomalo
Anomalo is an AI-powered enterprise data quality monitoring platform that automatically detects data issues across warehouses and lakes without manual rule configuration. The platform uses machine learning to monitor structured and unstructured datasets for enterprises like Block and Discover Financial.

Apheris
Apheris provides governed, privacy-preserving data access and collaboration for AI and analytics across sensitive datasets.

Ayar Labs
Ayar Labs builds optical I/O and in-package photonics technology to reduce data-movement bottlenecks in large-scale AI and high-performance computing systems.

Bindwell
Bindwell is an AI-powered pesticide discovery company that uses machine learning models 4x faster than DeepMind's AlphaFold to screen billions of molecules and design safer, more effective crop protection products. Unlike traditional agtech software companies, Bindwell develops and licenses complete proprietary pesticide molecules to major agrochemical companies. Founded by teen entrepreneurs Tyler Rose and Navvye Anand through Y Combinator's W25 batch, the company is backed by General Catalyst and Paul Graham.

Biostate AI
A scalable biological data collection service providing multi-omics data for research.

Bronto
Modern logging and observability platform for AI applications and engineering teams, offering fast log ingestion, search, and alerting with a columnar storage architecture.

Buster
Buster is an open-source AI-native analytics platform that gives teams AI data analysts and engineers for their data stack.

Byteport
Byteport builds global upload acceleration infrastructure for 1GB-100TB files using DART (Dynamic Accelerated Record Transfer), a proprietary UDP-based protocol that is typically 10x faster than TCP. The platform serves AI clusters, robotics, satellite networks, drone fleets, and defense applications, with SDKs across Python, Java, .NET, C++, Node.js, iOS, and Android. Backed by Y Combinator Winter 2026.

Conduktor
Conduktor builds an enterprise data management platform for Apache Kafka, giving teams a unifying layer to monitor, secure, and govern real-time data streaming. It provides security controls, monitoring, and governance to help regulated enterprises scale their Kafka usage. The company is headquartered in London with US expansion underway.

Coralogix
Coralogix is a full-stack observability and security platform that unifies logs, metrics, traces, and security data using a streaming ('Streama') engine and customer-owned storage, purpose-built for AI-scale telemetry. It serves 5,000+ customers including IBM and JFrog. The Israel-founded company raised a $200M Series F in June 2026 at a $1.6B valuation, bringing total funding to $550M.

DataLane
DataLane is a New York-based data platform building an identity graph for every local business in America — a 'LinkedIn for the offline economy.' Its AI autonomously extracts, deduplicates, verifies, and connects data from hundreds of sources into structured business intelligence, delivered into systems like Salesforce and Snowflake. Customers include companies such as DoorDash, Square, and Paychex that need to find and verify real local business owners.

Definite
Definite combines a cloud data warehouse, metrics layer, notebooks, dashboards, and AI assistant workflows into an all-in-one analytics platform for faster self-serve analysis.

DualBird
DualBird provides a cloud-native hardware-software data and AI infrastructure engine that delivers 10-100x faster performance and 50-90% lower costs through FPGA-based acceleration.

EquiLibre Technologies
EquiLibre Technologies is a Prague-based AI and quantitative-finance company building reinforcement-learning-based autonomous trading agents for global financial markets. Founded by former DeepMind researchers (creators of DeepStack poker AI), it applies game theory and large-scale AI to algorithmic trading. It closed a Series A led by Creandum at a $500M+ valuation in mid-2026.

Espresso AI
Espresso AI uses generative AI and machine learning to automatically optimize SQL queries and reduce cloud compute costs by up to 70-80% for Snowflake data warehouse users. The platform integrates with existing data warehouse setups to analyze and optimize queries in real time using NLP, program synthesis, and reinforcement learning.

Evolvent AI
Evolvent AI builds data infrastructure for self-evolving AI agent systems, producing high-quality RL/SFT post-training data and simulated environments for coding, SWE, terminal, and long-horizon agent tasks. Its platform lets a population of agents inside an organization co-evolve over time by accumulating skills and shared memory across trajectories. The team, drawn from top universities and leading foundation-model groups, sells post-training data and agent evaluation infrastructure to frontier model labs.

Flatfile
AI-assisted data exchange platform that helps teams collect, map, validate, and transform messy customer data before it enters core systems.

Flow Computing
Flow Computing develops Parallel Processing Unit technology to accelerate next-generation CPUs for AI, edge, cloud, and parallel computing workloads.

Funnel
Funnel is a marketing intelligence platform that enables businesses to collect, prepare, and analyze data from various marketing channels. Founded in Stockholm, it provides pre-built connectors for data integration, marketing reporting, and advanced measurement capabilities.

Leen
Leen provides a unified API and data fabric for cybersecurity, helping security product teams, internal security teams, and MSPs integrate once with security tools and normalize vulnerability, endpoint, cloud, identity, compliance, and operational data.

Mage
Mage is an open-source, AI-native data pipeline platform that enables teams to build, run, and manage data pipelines for integrating and transforming data using Python, SQL, and R. Available as both open-source and enterprise versions, it provides real-time and batch pipeline orchestration.

MicroAGI
MicroAGI is a Berlin-based data research lab accelerating embodied AI by recording real-world human task execution through camera-wearing workers, generating structured egocentric video datasets for robotics and AI labs. The company runs active data-collection operations in NYC through its Shift app, targeting the fast-growing market for physical-world AI training data.

Noitom Robotics
Noitom Robotics develops embodied intelligence technology for humanoid robots, providing high-fidelity motion capture, teleoperation systems, and multimodal data acquisition platforms used to train robot perception and locomotion. The company positions itself as a data infrastructure provider — 'a robot company that doesn't build robots' — supplying scalable training data and tooling to robot manufacturers and embodied AI model teams. It spun out of Noitom Technology, a global leader in professional motion capture.

Polars
Polars is a blazingly fast DataFrames library written in Rust, offering Python, R, Node.js, and SQL bindings for efficient, multi-threaded data manipulation at scale.

PrismaX
PrismaX is a decentralized robotics service layer that sets standards for how robots, data, and human teleoperation are deployed to advance physical AI. The platform uses token incentives to crowdsource high-quality training data and operates a decentralized teleoperations network for robotics foundation model development. It combines AI, robotics, and Web3 to address data scarcity bottlenecks in autonomous robot development.

PromptQL
PromptQL is an agentic AI data-access and analysis platform built by the team behind Hasura, letting enterprises deploy LLM-powered agents that query, reason over, and act on their business data reliably. It combines a query planning layer with LLM workflows to deliver trustworthy automated data analysis for business operations. The company is San Francisco-based and operates as the AI-native successor product of Hasura.

Reworkd
Reworkd builds AI-powered web data extraction agents that automate the entire scraping pipeline at scale. Note: Company announced sunsetting as of February 2025.

Rune
Developer of the world's first DC data centers built exclusively for solar and wind power. Using proprietary chip design and smart controllers, Rune converts stranded and curtailed renewable energy into compute power at generation sites.

Shofo
Shofo is building the world's largest indexed video library — described as 'Common Crawl for Videos' — by continuously crawling and labeling public video content from social platforms like TikTok. It sells custom, multi-stage annotated datasets (object detection, activity recognition, semantic segmentation, reasoning annotations) to AI labs for pre-training and fine-tuning. Part of Y Combinator's W2026 batch.

Supper
AI-native agentic data platform that integrates with SaaS tools and data warehouses, cleanses and normalizes data, and enables self-serve insights through natural language.

Tilde
Tilde Research is an AI infrastructure company building efficient mixture-of-experts (MoE) training and inference systems. It develops novel architectures and hardware-aware algorithms to reduce the cost and improve the performance of large language model training.
FAQ
What is the Data Engineering tag page on Fast AI Startup Jobs?
It is a curated landing page that groups AI startup companies tagged with Data Engineering, plus links to their company profiles and available jobs.
How many Data Engineering companies are included?
This page currently lists 102 companies tagged with Data Engineering.
How many jobs are associated with Data Engineering companies?
The companies on this page currently account for 2280 listed jobs in our public dataset (subject to regular updates).
What roles are most common at Data Engineering companies?
Based on currently listed jobs for Data Engineering companies, the most common role groups are Engineering (1981), Other (692), Sales (512).
What funding stages are most common among Data Engineering companies?
Common funding stages on this Data Engineering page include Series A (28), Seed (27), Series B (17), Series C (9).
Where do the job links go?
Job links point to official company career pages or public job listings, not re-hosted application forms.
How often is this tag page refreshed?
Data is refreshed on a near-daily cadence as public company and job listings change.