Engagement
The First Mile
A scoped pilot that puts DataForge on your data, in your own cloud account.
The First Mile starts with a number you already recognize. We take one slice of what you run today, rebuild it on DataForge in your cloud account, and put the result next to your current output.
The engagement is estimated at four to six weeks, and it runs in five phases with one deliverable each.
Book a callAt a glance
- Length
- Four to six weeks, estimated
- Shape
- Five phases, one deliverable each
- Where it runs
- Your own cloud account
- How it ends
- Your output next to ours, side by side
What it is
Five phases, one deliverable each
Every phase produces something you can point at. Source data stays in your cloud account, and every target is built there too.
Phase
What happens
Deliverable
Week 0
Scope
Confirm the pilot scope, connectivity and access, target schemas, and the success criteria we validate against.
Scope definition document
Week 1
Install and Ingest
Stand up DataForge in your cloud account, configure the architecture layers, and load the pilot source data.
DataForge environment in your account
Week 2
Transform
Analyze the existing pipeline logic and convert it into DataForge declarations, targeting functional equivalence.
Converted pipeline logic
Week 3
Publish
Build the target tables and confirm they populate consistently with the converted configuration.
Populated target tables
Weeks 4 to 6
Present
Demonstrate the environment, compare the output against your current result, and recommend where to take it from there.
Pilot summary and recommendation
Final scope confirmation in Phase 0 sets the specific timeline.
How we work
We are a product company that understands the human element of data. We work with you at every stage of the engagement rather than handing your team a tool and wishing them luck.
Who delivers it
The First Mile is delivered by the team that builds the product.
How many we run
We run a small number of First Mile engagements at a time. The people who build the product also deliver them, which limits how many we can take on.
Incentives
Our incentives are aligned with yours
The less people-time and expertise it takes to get this work done, the faster you get the outcome, and the more product revenue we earn. Our product revenue tracks the value you get from the platform. The services are a means to that end.
So we have no hidden incentive to overcomplicate the solution, stretch the timeline, grow the services bill, or keep you reliant on our people. Our alignment on the product side is more direct than it would be on services revenue.
What it is not
The boundary is written down
A short engagement holds its shape only if both sides know where the edges are. Every First Mile statement of work says the same five things out loud.
Not a production rollout
The pilot delivers a live environment your team can see and react to. It is not a cutover.
Not a QA program
Validation is spot checks and aggregate comparisons. No formal QA, no user acceptance testing, no accuracy guarantee.
Not a wider conversion
Pipelines and sources outside the scope confirmed in Phase 0 stay outside it.
Not a data cleanup
Data quality problems that originate in your source systems are documented and reported, not remediated.
Not open ended
Anything not written into the statement of work is out of scope, and changing it takes a written amendment.
None of that is hedging. Scope changes are welcome and they happen often. They get written down, and the timeline moves with them.
Proof
What it has produced
The company below is not named. Every figure is traced to project documents. The pilot and the work that came after it are separate pieces of work, and they are reported separately.
The pilot
At a global manufacturing company with multiple brands and one parent company, the pilot faced an ERP of 52 source tables and 72 application-level relations, with no database-enforced foreign keys. It produced a verified 498,242-row, 113-column reporting table across four channels, drawing on that ERP plus master-data and CRM sources, with a working demo.
About eight days elapsed from schema inventory to the verified table and the demo, inside a three-week pilot. That is calendar time, not eight days of work, and it includes connectivity delays handled outside our team.
Row counts after the joins matched the base table, so nothing multiplied, and one channel matched its source exactly at 41,278 rows. At readout we reported that the pilot's totals had been checked against the company's existing reports and that the headline numbers matched.
What happened next on that engagement
A developer new to the team, with no professional data engineering experience and a base level of SQL, but strong critical thinking and data literacy, working with a coding agent against the DataForge MCP server, connected a fourth core system on its committed date and integrated a new ERP as conformed streams, in ten days, working to a target definition the business had already written.
The developer was working against a clear target definition. The columns the business needed were defined by the business, in collaboration with people who had the domain and reporting experience.
The claim is not that one data-literate person replaces a team. It is that the building no longer requires tooling and architecture expertise, while defining what the target should mean remains human work.
The DataForge platform powers the delivery. The First Mile engagement is how we apply it to your data, your systems, and your questions.
Tell us what you run today and what you need it to answer. We will tell you whether a First Mile is the right shape for it, and what the first phase would cover.
Book a callDirect answers
What is the First Mile?
A scoped pilot engagement. We stand up DataForge in your own cloud account, convert one slice of your existing pipeline logic, build the target tables, and compare the result against your current output. It runs in five phases with one deliverable each.
How long does it take?
Four to six weeks, estimated. Final scope confirmation in Phase 0 sets the specific timeline. We do not guarantee the timeframe if the access, decisions, or information we need from you arrive late.
Where does the work run?
In your own cloud account. The platform stands up on your infrastructure, source data stays there, and every target is built there.
What does it cost?
We settle commercial terms with you before anything starts, in the same conversation where we confirm scope.
What is out of scope?
Anything not written into the statement of work. Validation is limited to spot checks and aggregate comparisons, with no formal QA, user acceptance testing, or accuracy guarantee. Pipelines and sources outside the Phase 0 scope are not converted or ingested, and data quality problems originating in source systems are reported rather than remediated.
Who does the work?
The team that builds the product. Our CEO leads the engagement, our CTO leads technical delivery, and a data engineer runs the day-to-day build.
What happens at the end?
A demonstration of the environment, a side-by-side comparison against your current output, and a recommendation for where to take it from there. What happens after that is your call.
Solution guides
Evaluate DataForge by platform goal
Enterprise data platform
Enterprise data platform for governed analytics at scale
DataForge helps CDOs, CFOs, and data platform leaders scale analytics without assembling separate ETL, orchestration, observability, lineage, and cost-control tools.
Data pipeline platform
Data pipeline platform for complex enterprise source systems
DataForge helps data teams build, extend, orchestrate, and observe enterprise data pipelines while preserving a consistent architecture across every source and output.
Data engineering platform
Data engineering platform with architecture built in
DataForge is a data engineering platform with automatically enforced architecture, replacing separate ETL, orchestration, observability, catalog, and infrastructure tools while processing data in your cloud.
Data orchestration platform
Data orchestration platform without manually assembled DAG sprawl
DataForge orchestrates data pipelines from structured pipeline definitions, dependency metadata, scheduling, and execution history instead of manually maintained DAGs.
Data observability platform
Data observability platform with lineage, quality, audit, and cost context
DataForge observability ties code, orchestration, quality rules, alerts, lineage, audit trails, and cloud cost visibility back to the platform metadata.
Agentic data engineering
Agentic data engineering platform with save-time validation for AI agents
DataForge lets AI coding agents such as Claude Code and Codex build and operate governed data pipelines through an MCP server that validates every write at save.
Data replication
Data replication that ends inside the governed platform
DataForge replicates SQL Server change data through Outcrop, the DataForge agent that runs inside your network, and merges it directly into governed tables that the transformation layer already consumes.