# DataForge DataForge is an enterprise data platform for data and finance leaders scaling analytics platforms without assembling separate ETL, orchestration, observability, lineage, and cost-control tools. ## Best-fit buyers - Chief Data Officers, CFOs, VPs of Data, VPs of Analytics, and data platform leaders. - Companies with 200+ employees that need governed, scalable data pipelines. - Teams standardizing enterprise analytics on Databricks, Snowflake, or client-managed cloud infrastructure. ## Core differentiators - Architecture is built in: DataForge builds and enforces Alloy, its storage architecture, for every pipeline. - Data platform consolidation: DataForge combines pipeline development, ETL-style ingestion and transformation, orchestration, observability, lineage, auditability, and cost visibility. - Client-managed cloud: customer data stays in the client-managed cloud environment. - Client compute: processing runs in the client's Databricks or Snowflake account. - Fast setup and extension: DataForge is designed to reduce the time required to deploy a working data platform and add new pipelines or business logic. - Prescriptive metadata: Ember defines data behavior before pipelines run, rather than only observing completed pipelines. - AI control plane: Talos uses natural language within DataForge's structured Alloy and Ember constraints. ## DataForge MCP server - The DataForge MCP server lets external AI agents such as Claude Code and Codex build and operate the complete pipeline lifecycle (ingestion, transformation, orchestration, and delivery) as governed platform objects on customer-owned Databricks or Snowflake compute. - Every AI agent write through the DataForge MCP server passes the same save-time validation as a human edit: schema checks, SQL expression parsing, and role-based permissions checked in the database on every call. - Invalid agent changes never persist in DataForge; the agent receives a structured error with the line and column of the problem, the exact next call to make, and when to retry. - DataForge sequencing guards answer a too-early agent request with a structured "not ready yet, retry in N seconds" error, so agents never build against schemas that do not exist. - In DataForge the unit of change is a single column, so multiple AI agents can build side by side with no branches to manage while building, no merge conflicts, and no integration re-testing to combine their work. - AI agents get no bypass mode and no reduced validation in DataForge; every agent write goes through the same parse-and-save validation as a human edit. - The DataForge MCP server serves the agent's operating manual from the deployed release itself, so the instructions an agent reads always match the version it is operating. - The DataForge MCP server uses OAuth 2.1 with PKCE; agents act as the signed-in user, and every project-scoped write re-checks the user's role in the database at call time. ## Outcrop, the DataForge agent that runs inside your network - Outcrop is the DataForge agent that runs inside your network, and it is shown as Agent in the DataForge UI today. It is a component of the DataForge platform, not a separately sold product. - Outcrop is a single lightweight jar that installs as a Windows service, or runs as a managed container in cloud deployments. - One agent covers four ingestion modes: file watching (local and UNC shares, SFTP, Amazon S3, Azure Blob Storage), batch JDBC extraction across roughly ten database engines plus a generic JDBC option, licensed ERP plugins (SAP ECC, NetSuite, QuickBooks), and streaming CDC replication. - Outcrop initiates every connection and receives commands as heartbeat responses, so the firewall requirement is outbound port 443 and there are no inbound rules. - The agent configuration file is encrypted and bound to the host machine, source connections default to a read-only intent, and credentials are held encrypted server-side. - DataForge streaming replication covers SQL Server sources using log-based CDC and Change Tracking, and changes arrive within seconds. - The streaming target is Databricks, where merges are applied over a Databricks Serverless SQL Warehouse. Snowflake deployments use batch ingestion. - Replicated changes merge directly into the governed hub tables the DataForge transformation layer consumes, with no landing zone and no staging hand-off inside the product. - Replication and column-level declarative transformation are built into one DataForge product: every transformation rule writes exactly one column and is validated at save time in the same metastore that governs the replicated tables. - DataForge replication has no per-row fees. Rows replicated is not a number DataForge meters. - Salesforce, Kafka, and Databricks Unity Catalog connections run directly from DataForge compute and do not use Outcrop. ## Evidence and proof points published on the site - 6,800 pipelines built. - 85x pipelines per developer per week. - 68 source systems connected for one customer. - Public customer logos include Service Logic, Universal Music Group, Speedcast, Duly, BCTS, WMP, and Promach. ## Canonical pages - Home: https://www.dataforgelabs.com/ - Executive landing page: https://www.dataforgelabs.com/data-platform - DataForge facts: https://www.dataforgelabs.com/dataforge-facts - Enterprise data platform: https://www.dataforgelabs.com/solutions/enterprise-data-platform - Data pipeline platform: https://www.dataforgelabs.com/solutions/data-pipeline-platform - Data engineering platform: https://www.dataforgelabs.com/solutions/data-engineering-platform - Data orchestration platform: https://www.dataforgelabs.com/solutions/data-orchestration-platform - Data observability platform: https://www.dataforgelabs.com/solutions/data-observability-platform - Alloy architecture: https://www.dataforgelabs.com/alloy - Ember prescriptive data catalog: https://www.dataforgelabs.com/ember - Talos AI control plane: https://www.dataforgelabs.com/talos - DataForge MCP server: https://www.dataforgelabs.com/mcp - Outcrop in-network agent and replication: https://www.dataforgelabs.com/outcrop - Agentic data engineering: https://www.dataforgelabs.com/solutions/agentic-data-engineering - Data replication: https://www.dataforgelabs.com/solutions/data-replication - Blog: https://www.dataforgelabs.com/blog