lakeFS logo

AI visibility report

lakeFS ranks #1 in AI Data Curation and Dataset Versioning AI search.

Outside the top three on 3 of the 25 prompts buyers actually ask.

Encord is cited on 3 of those losses.

25 prompts
5 platforms
Updated Jul 11, 2026 - refreshed weekly
Track lakeFS daily

Free trial. Setup comes pre-filled for lakeFS.

Track lakeFS across these prompts daily.

Start free trial
38percent
Presence Rate
Weak presence

Best among 7 vendors · still absent from 62.4% of tracked prompt responses

Top-3 citations across 125 prompt × platform pairs

+0.60
Sentiment
-1.00.0+1.0
Very positive
#1of 7

Peer Ranking

#1#7
Top tierin AI Data Curation and Dataset Versioning

Key Metrics

Presence Rate37.6%
Share of Voice70.2%
Avg Position#8.3
Docs Presence15.2%
Blog Presence26.4%
Brand Mentions56.0%

Platform Breakdown

ChatGPT
72%18/25 prompts
Google AI Mode
36%9/25 prompts
Gemini Search
36%9/25 prompts
Perplexity
36%9/25 prompts
Bing Copilot
8%2/25 prompts

Leader, with room to expand. lakeFS leads this category on presence and share of voice, but appears in only 37.6% of tracked prompt responses. The priority is defending current wins while expanding absolute coverage.

Where lakeFS is losing

Prompts where competitors are visible and lakeFS is not.

These prompt-level losses are the first prompts to track and repair.

Where lakeFS is winning5

  • What data curation platforms work well alongside a feature store and a model registry for a fully lineage-tracked ML pipeline?

    Avg # 1.0 · 1 platform

  • What's the quickest way for an ML engineer to set up reproducible dataset snapshots without migrating away from existing cloud object storage?

    Avg # 1.0 · 3 platforms

  • Which data curation tools handle mixed-modality datasets — images, text, and structured metadata — in a single versioned artifact?

    Avg # 1.7 · 3 platforms

  • Which data curation platforms integrate with workflow orchestration tools so dataset preprocessing and versioning steps run as part of an automated ML pipeline?

    Avg # 2.0 · 3 platforms

  • Which dataset versioning tools integrate best with experiment tracking platforms so training runs automatically link to the exact dataset version used?

    Avg # 2.0 · 2 platforms

Where lakeFS is losing3

  • Which AI dataset management tools have native connectors to annotation and labeling services so curated slices can be sent for labeling without manual export?

    Competitors on 2 platforms

    Track this prompt
  • I'm evaluating data curation platforms for a computer vision team of 5 — which ones have the fastest path from raw images to a labeled, versioned dataset?

    Competitors on 1 platform

    Track this prompt
  • Which AI data curation platforms offer the best visual dataset explorer so non-engineers on a labeling team can review samples without writing code?

    Competitors on 1 platform

    Track this prompt

Track lakeFS daily before the next report refresh.

Track these gaps
Research dossierCapabilities, use cases, sources, reviews, pricing, and FAQ

Overview

lakeFS, developed by Treeverse, is an open-source data version control system that applies Git-like operations—branching, committing, merging, and reverting—to data lakes stored in object storage (Amazon S3, Azure Blob Storage, Google Cloud Storage). Founded in 2020 by Oz Katz and Dr. Einat Orr, lakeFS operates as a metadata layer atop existing storage without moving or duplicating data. It enables data and AI/ML teams to create isolated environments for testing, ensure reproducibility of model training, enforce data quality via CI/CD hooks, roll back from data incidents, and maintain full data lineage and audit trails. Available as a free open-source edition and a commercial Enterprise tier, lakeFS is used by organizations including Arm, Netflix, Volvo, Lockheed Martin, NASA, and the U.S. Department of Energy. In November 2025, lakeFS acquired the DVC open-source project from Iterative.ai, extending its reach from enterprise to individual practitioners.

lakeFS is an open-source and enterprise data version control platform that transforms object storage into Git-like repositories, enabling data and AI teams to branch, commit, merge, and roll back datasets at petabyte scale without copying data. Built by Treeverse and backed by $43M in funding, it supports reproducible ML workflows, data quality enforcement, and governance across multi-cloud and on-premises data lakes, with deep integration across the modern data and AI tooling stack.

Key Facts

Founded
2020
HQ
New York, NY, USA
Founders
Oz Katz, Einat Orr
Employees
11-50
Funding
$43M
Customers
>90,000 organizations
Status
Private

Target users

Data engineers building and maintaining large-scale data lakes and ETL pipelinesML/AI engineers and MLOps teams managing model training data and experimentsData scientists requiring reproducible, versioned datasets for researchDataOps and platform engineering teams at data-intensive enterprisesData governance and compliance officers in regulated industries (defense, healthcare, energy)CTOs and engineering leaders overseeing AI/ML infrastructure at scale

Key Capabilities10

  • Git-like branching, committing, merging, and reverting for object storage data lakes at petabyte scale
  • Zero-copy isolated dev/test environments via branches without data duplication
  • Atomic commits and instant rollback for data pipeline error recovery
  • Data CI/CD via configurable pre- and post-commit hooks for quality gates
  • Full data lineage tracking and built-in audit trail for governance and compliance
  • S3-compatible API enabling seamless integration with existing tools and frameworks
  • Role-based access control (RBAC), SSO, and SCIM support (Enterprise tier)
  • Iceberg REST Catalog support and format-agnostic versioning (structured and unstructured)
  • lakeFS Mount for local filesystem-style access to remote data without full downloads
  • Transactional mirroring and multi-storage backend support (Enterprise tier)

Key Use Cases8

  • Reproducible ML/AI model training with versioned, immutable dataset snapshots
  • Isolated ETL testing on production data without copying or risking production state
  • Data quality enforcement via Write-Audit-Publish pipeline patterns
  • Instant rollback and recovery from bad data incidents in production data lakes
  • ML experiment tracking tied to specific data versions for auditability
  • Data governance and compliance for regulated industries (FDA, DOE, defense)
  • Multi-team collaboration on shared data lakes with branch-level isolation
  • Managing petabyte-scale multimodal AI training data lifecycle across cloud environments

lakeFS customer outcomes

Enigma

80% reduction in testing time

Enigma adopted lakeFS data branching and within days of migration had reduced testing time on two different data pipeline projects, with CTO Ryan Green crediting data branching for improved product velocity.

Ellips

6 models launched per 2-week sprint with half the team (vs. 2–3 models with full team prior)

Ellips's ML engineering team implemented lakeFS and dramatically accelerated model deployment cadence, launching more models per sprint with a smaller team than previously possible.

Arm

Arm implemented lakeFS for automated data cleaning, version control, and governance across distributed teams, resulting in faster go-to-market, reduced storage costs, improved development velocity, and stronger data governance.

Paige AI

Paige AI used lakeFS alongside dbt to enable reproducible ML experiments, increase data team productivity, and satisfy FDA compliance requirements for AI-powered cancer diagnostics.

Recent Trend

Visibility+12.0 pts
Avg position+0.69
Sentiment+0.17

How AI describes lakeFS3

The easiest Python‑native libraries for versioning multimodal training data directly on top of an existing object‑storage bucket are lakeFS, DVC, Metaxy, and DataChain.

Which Python-native dataset management libraries make it easiest to start versioning multimodal training data from an existing object storage bucket?

bing-copilot-searchDirect lakeFS mention
Today, that effectively means lakeFS, Pachyderm, and Delta Lake / Apache Iceberg when used in a lakehouse with high‑throughput readers.

Which dataset versioning tools handle petabyte-scale training datasets without bottlenecking the data loading pipeline during distributed training?

bing-copilot-searchDirect lakeFS mention
The strongest multi‑cloud dataset‑versioning tools today are lakeFS, DVC, Pachyderm, ActiveLoop Deep Lake, and Oxen.ai — each with different strengths depending on whether you need Git‑like branching, pipeline lineage, multimodal...

Looking for a dataset versioning tool that works with multiple cloud object storage providers to avoid lock-in — what are the best options?

bing-copilot-searchDirect lakeFS mention

Alternatives in AI Data Curation and Dataset Versioning6

lakeFS positions itself as the enterprise-grade 'control plane for AI-ready data,' differentiating through Git-like branching and versioning applied at petabyte scale to object storage (S3, GCS, Azure Blob).

  • Unlike annotation- or labeling-focused tools in the AI data curation space (Encord, Roboflow, Voxel51), lakeFS operates at the data infrastructure layer, providing reproducibility, lineage, and governance for data lakes underpinning AI/ML pipelines.
  • Its November 2025 acquisition of DVC from Iterative.ai extended market coverage from enterprise data engineering teams down to individual data scientists.
  • It was named a Representative Vendor in the 2025 Gartner Market Guide for DataOps Tools, and is one of the few open-source-core data version control systems with a commercial enterprise tier at this scale.
View category comparison hub

Reviews

Praised

  • Familiar Git-like UX for data engineers and developers
  • Zero-copy branching with no data duplication overhead
  • S3 API compatibility requiring no changes to existing tools
  • Fast quickstart and easy local sandbox setup
  • Format-agnostic versioning (Parquet, images, video, JSON, CSV)
  • Active open-source community, Slack support, and sample notebooks
  • Atomic commits that prevent partial or inconsistent data states
  • Broad integration coverage across modern data and ML stack

Criticized

  • Enterprise features gated behind contact-sales pricing with no public tiers
  • Self-hosted open-source deployment can be complex for smaller teams
  • Limited publicly verifiable third-party reviews on enterprise platforms
  • Architecture tightly coupled to S3-compatible object storage
  • Governance features (RBAC, SSO, audit logs) only available in Enterprise tier
  • Historically less accessible for individual data scientists on small datasets

Formal review scores are not publicly available on major enterprise review platforms (G2, Gartner Peer Insights, TrustRadius) as of early 2026, reflecting lakeFS's open-source heritage and early commercial maturity. Community sentiment from GitHub (5.3k stars, 446 forks) and Slack is strongly positive, with practitioners praising Git-like UX, zero-copy branching, and broad S3 compatibility. Independent technical evaluations (e.g., Data Minded, 2021) historically noted deployment complexity and overhead for smaller teams, though the product has evolved significantly since.

Pricing

lakeFS offers two tiers: Open Source (free forever, self-hosted, includes core data version control, branching, merging, hooks, garbage collection, and S3 API compatibility) and Enterprise (unlimited seats, contact sales for pricing, adds RBAC, SSO, SCIM, IAM Roles, lakeFS Mount, Audit Logs, Transactional Mirroring, Iceberg REST Catalog, Metadata Search, multi-storage backend support, simplified garbage collection, SOC2 certification, and a support SLA). Cloud-hosted access ('Try lakeFS') is available. No public per-seat or consumption-based Enterprise pricing is disclosed.

Limitations

  • Enterprise-tier features—including RBAC, SSO, SCIM, audit logs, lakeFS Mount, Iceberg REST Catalog, transactional mirroring, and SLA-backed support—require contacting sales with no published pricing. lakeFS is architecturally tied to S3-compatible object storage and is less suitable for non-object-storage data environments.
  • Structured review presence on major enterprise platforms (G2, Gartner Peer Insights, TrustRadius) is minimal, limiting third-party validation for procurement teams.
  • Self-hosting the open-source edition at scale may require significant DevOps expertise.
  • The platform has historically been less suited for lightweight individual data science workflows, a gap partially addressed by the November 2025 acquisition of DVC.

Frequently asked questions

Topic coverageCoverage by buyer topic

Topic Coverage

Capability4/5DevEx4/5Integrations &Ecosystem4/5Performance &Reliability4/5Setup & First Run5/5

Prompt-Level Results

Brand citedCompetitor citedNot cited
PromptBing CopilotChatGPTGoogle AI ModeGemini SearchPerplexity
Capability4/5 cited (80%)

What are the best platforms for embedding-based deduplication and near-duplicate detection across a multimodal training corpus at scale?

Which dataset versioning tools support branching and merging semantics similar to source control for managing parallel data experiments?

Which data curation tools handle mixed-modality datasets — images, text, and structured metadata — in a single versioned artifact?

Which dataset versioning platforms let ML teams tag, lineage-track, and roll back to any historical dataset state used for a production model training run?

I'm evaluating AI dataset management tools — which ones support automated data quality checks and slice-level statistics for model evaluation sets?

Developer Experience4/5 cited (80%)

Which dataset versioning platforms have the best Python SDK for iterating over large image datasets without loading everything into memory?

Which dataset management tools make it easiest for ML engineers to query, filter, and tag unstructured data using embedding-based similarity search?

Which AI data curation platforms offer the best visual dataset explorer so non-engineers on a labeling team can review samples without writing code?

Looking for a dataset versioning tool with a great notebook-friendly workflow — what are the best options for teams that live in Jupyter?

What data curation tools do ML platform teams typically use to give model trainers a clean, reproducible slice of a dataset without raw storage access?

Integrations & Ecosystem4/5 cited (80%)

What data curation platforms work well alongside a feature store and a model registry for a fully lineage-tracked ML pipeline?

Which data curation platforms integrate with workflow orchestration tools so dataset preprocessing and versioning steps run as part of an automated ML pipeline?

Looking for a dataset versioning tool that works with multiple cloud object storage providers to avoid lock-in — what are the best options?

Which dataset versioning tools integrate best with experiment tracking platforms so training runs automatically link to the exact dataset version used?

Which AI dataset management tools have native connectors to annotation and labeling services so curated slices can be sent for labeling without manual export?

Performance & Reliability4/5 cited (80%)

Which AI dataset platforms have the best performance for querying embedding indexes across tens of millions of vectors in a curation workflow?

What are the best AI data curation platforms for streaming random-access reads from large image datasets stored in object storage with low latency?

Which dataset versioning tools handle petabyte-scale training datasets without bottlenecking the data loading pipeline during distributed training?

Which dataset management tools are production-proven for enterprise ML teams managing hundreds of dataset versions without storage cost spiraling?

Looking for a data versioning layer over object storage that handles concurrent writes from multiple experiment runs without corruption — what are my options?

Setup & First Run5/5 cited (100%)

Which Python-native dataset management libraries make it easiest to start versioning multimodal training data from an existing object storage bucket?

What are the best dataset versioning tools for a small ML team to get started with object storage without a complex infrastructure setup?

I'm evaluating data curation platforms for a computer vision team of 5 — which ones have the fastest path from raw images to a labeled, versioned dataset?

Which dataset versioning tools have the smoothest onboarding for teams migrating off a manual folder-based data management system?

What's the quickest way for an ML engineer to set up reproducible dataset snapshots without migrating away from existing cloud object storage?

Turn this matrix into daily prompt monitoring.

Track prompt changes

Vertical Ranking

#BrandPres.SoVDocsBlogMent.PosSentiment
1lakeFS37.6%70.2%15.2%26.4%56.0%#8.3+0.60
2Encord10.4%12.3%0.0%10.4%16.0%#6.8+0.53
3Voxel513.2%6.7%2.4%0.8%7.2%#9.6+0.70
4Activeloop3.2%5.2%0.0%0.8%12.8%#10.0+0.60
5Roboflow2.4%4.0%0.0%2.4%5.6%#13.7+0.43
6DataChain0.8%1.2%0.8%0.0%2.4%#7.0+0.80
7Nomic AI0.8%0.4%0.0%0.0%0.0%#10.0+0.30

Turn this into your team dashboard

Sign up to unlock project-level analytics, daily tracking, actionable insights, custom prompt configurations, adoption tracking, AI traffic analytics and more.

Free trial. Setup comes pre-filled from this report.

Get started free