Engagement

The First Mile

A scoped pilot that puts DataForge on your data, in your own cloud account.

The First Mile starts with a number you already recognize. We take one slice of what you run today, rebuild it on DataForge in your cloud account, and put the result next to your current output.

The engagement is estimated at four to six weeks, and it runs in five phases with one deliverable each.

Book a call

At a glance

Length
Four to six weeks, estimated
Shape
Five phases, one deliverable each
Where it runs
Your own cloud account
How it ends
Your output next to ours, side by side

What it is

Five phases, one deliverable each

Every phase produces something you can point at. Source data stays in your cloud account, and every target is built there too.

Week 0

Scope

Confirm the pilot scope, connectivity and access, target schemas, and the success criteria we validate against.

Scope definition document

Week 1

Install and Ingest

Stand up DataForge in your cloud account, configure the architecture layers, and load the pilot source data.

DataForge environment in your account

Week 2

Transform

Analyze the existing pipeline logic and convert it into DataForge declarations, targeting functional equivalence.

Converted pipeline logic

Week 3

Publish

Build the target tables and confirm they populate consistently with the converted configuration.

Populated target tables

Weeks 4 to 6

Present

Demonstrate the environment, compare the output against your current result, and recommend where to take it from there.

Pilot summary and recommendation

Final scope confirmation in Phase 0 sets the specific timeline.

How we work

We are a product company that understands the human element of data. We work with you at every stage of the engagement rather than handing your team a tool and wishing them luck.

Who delivers it

The First Mile is delivered by the team that builds the product.

How many we run

We run a small number of First Mile engagements at a time. The people who build the product also deliver them, which limits how many we can take on.

Incentives

Our incentives are aligned with yours

The less people-time and expertise it takes to get this work done, the faster you get the outcome, and the more product revenue we earn. Our product revenue tracks the value you get from the platform. The services are a means to that end.

So we have no hidden incentive to overcomplicate the solution, stretch the timeline, grow the services bill, or keep you reliant on our people. Our alignment on the product side is more direct than it would be on services revenue.

What it is not

The boundary is written down

A short engagement holds its shape only if both sides know where the edges are. Every First Mile statement of work says the same five things out loud.

Not a production rollout

The pilot delivers a live environment your team can see and react to. It is not a cutover.

Not a QA program

Validation is spot checks and aggregate comparisons. No formal QA, no user acceptance testing, no accuracy guarantee.

Not a wider conversion

Pipelines and sources outside the scope confirmed in Phase 0 stay outside it.

Not a data cleanup

Data quality problems that originate in your source systems are documented and reported, not remediated.

Not open ended

Anything not written into the statement of work is out of scope, and changing it takes a written amendment.

None of that is hedging. Scope changes are welcome and they happen often. They get written down, and the timeline moves with them.

Proof

What it has produced

The company below is not named. Every figure is traced to project documents. The pilot and the work that came after it are separate pieces of work, and they are reported separately.

The pilot

At a global manufacturing company with multiple brands and one parent company, the pilot faced an ERP of 52 source tables and 72 application-level relations, with no database-enforced foreign keys. It produced a verified 498,242-row, 113-column reporting table across four channels, drawing on that ERP plus master-data and CRM sources, with a working demo.

About eight days elapsed from schema inventory to the verified table and the demo, inside a three-week pilot. That is calendar time, not eight days of work, and it includes connectivity delays handled outside our team.

Row counts after the joins matched the base table, so nothing multiplied, and one channel matched its source exactly at 41,278 rows. At readout we reported that the pilot's totals had been checked against the company's existing reports and that the headline numbers matched.

What happened next on that engagement

A developer new to the team, with no professional data engineering experience and a base level of SQL, but strong critical thinking and data literacy, working with a coding agent against the DataForge MCP server, connected a fourth core system on its committed date and integrated a new ERP as conformed streams, in ten days, working to a target definition the business had already written.

The developer was working against a clear target definition. The columns the business needed were defined by the business, in collaboration with people who had the domain and reporting experience.

The claim is not that one data-literate person replaces a team. It is that the building no longer requires tooling and architecture expertise, while defining what the target should mean remains human work.

The DataForge platform powers the delivery. The First Mile engagement is how we apply it to your data, your systems, and your questions.

Tell us what you run today and what you need it to answer. We will tell you whether a First Mile is the right shape for it, and what the first phase would cover.

Book a call

Direct answers

What is the First Mile?

A scoped pilot engagement. We stand up DataForge in your own cloud account, convert one slice of your existing pipeline logic, build the target tables, and compare the result against your current output. It runs in five phases with one deliverable each.

How long does it take?

Four to six weeks, estimated. Final scope confirmation in Phase 0 sets the specific timeline. We do not guarantee the timeframe if the access, decisions, or information we need from you arrive late.

Where does the work run?

In your own cloud account. The platform stands up on your infrastructure, source data stays there, and every target is built there.

What does it cost?

We settle commercial terms with you before anything starts, in the same conversation where we confirm scope.

What is out of scope?

Anything not written into the statement of work. Validation is limited to spot checks and aggregate comparisons, with no formal QA, user acceptance testing, or accuracy guarantee. Pipelines and sources outside the Phase 0 scope are not converted or ingested, and data quality problems originating in source systems are reported rather than remediated.

Who does the work?

The team that builds the product. Our CEO leads the engagement, our CTO leads technical delivery, and a data engineer runs the day-to-day build.

What happens at the end?

A demonstration of the environment, a side-by-side comparison against your current output, and a recommendation for where to take it from there. What happens after that is your call.

Solution guides

Evaluate DataForge by platform goal

View DataForge facts

Enterprise data platform

Enterprise data platform for governed analytics at scale

DataForge helps CDOs, CFOs, and data platform leaders scale analytics without assembling separate ETL, orchestration, observability, lineage, and cost-control tools.

Data pipeline platform

Data pipeline platform for complex enterprise source systems

DataForge helps data teams build, extend, orchestrate, and observe enterprise data pipelines while preserving a consistent architecture across every source and output.

Data engineering platform

Data engineering platform with architecture built in

DataForge is a data engineering platform with automatically enforced architecture, replacing separate ETL, orchestration, observability, catalog, and infrastructure tools while processing data in your cloud.

Data orchestration platform

Data orchestration platform without manually assembled DAG sprawl

DataForge orchestrates data pipelines from structured pipeline definitions, dependency metadata, scheduling, and execution history instead of manually maintained DAGs.

Data observability platform

Data observability platform with lineage, quality, audit, and cost context

DataForge observability ties code, orchestration, quality rules, alerts, lineage, audit trails, and cloud cost visibility back to the platform metadata.

Agentic data engineering

Agentic data engineering platform with save-time validation for AI agents

DataForge lets AI coding agents such as Claude Code and Codex build and operate governed data pipelines through an MCP server that validates every write at save.

Data replication

Data replication that ends inside the governed platform

DataForge replicates SQL Server change data through Outcrop, the DataForge agent that runs inside your network, and merges it directly into governed tables that the transformation layer already consumes.