SERVICES
On 23 September 2026, Bianca Stratulat made her Big Data LDN debut in the Data for Good Theatre with a session titled Saying Yes to Data for Good. Her talk drew on a recent Unifeye engagement: a production-ready Databricks platform for a public sector organisation, delivered in just eight weeks.
A 30-minute conference slot is enough to share the headlines, but not the full story. The questions Bianca received during and after the session, particularly on governance, cost and capability building, deserved more space. In this article she expands on the talk and goes deeper into the decisions, trade-offs and lessons behind the project.
Saying yes, with conditions
The session was never about saying yes to everything. It was about understanding what needs to be true before you can confidently say yes to something ambitious.
The starting point will be familiar to many organisations, particularly in the public sector: a legacy data estate, growing reporting demands, constrained budgets and critical knowledge held by very few people. Hundreds of reports. Bespoke SQL logic. Replication processes. And rising expectations around governance, security and what the organisation could achieve with its data.
The easy answer was “this will take a long time.” Instead, we asked a different question: what would it take to make meaningful change in eight weeks?
We didn’t mean a proof of concept, a set of notebooks someone would need to productionise later, or an architecture diagram of a hypothetical future. We meant a real, production-ready Databricks platform with governance, multiple environments, automated deployment, ingestion, transformation, reporting, monitoring and knowledge transfer. Above all, it had to be a platform the client’s team could own once we left.
That final requirement turned out to be the most important lesson of the engagement.
The starting point: technical debt isn’t always code
The organisation wasn’t starting from nothing. Its systems worked. They included a Progress database and Capita OpenHousing, Pro2 replication into SQL Server, stored procedures holding vital business logic, and approximately 550 SSRS reports serving teams across the business.
The problem wasn’t that the technology was old. Legacy technology isn’t automatically bad technology, and replacing something stable, understood and valuable simply because something newer exists isn’t transformation.
The real challenge was the operating model that had grown around it:
• Critical knowledge sat with a single person, supported by undocumented procedures.
• Bespoke stored procedures and increasingly complex SSRS queries were consuming SQL Server compute.
• Replication and reporting logic created dependencies that were becoming harder to change safely.
This is a form of technical debt that rarely appears on an architecture diagram: knowledge concentration. If only one or two people understand how something works, the organisation doesn’t truly own that capability. Those individuals do.
Modernisation therefore couldn’t mean moving the same dependency onto a newer platform. We had to modernise the technology and the ownership model together.
Eight weeks changes how you think
The constraints were clear: eight weeks, three environments, four source systems, two priority reports and a deliberately small delivery team.
It would be easy to assume something had to give. Nothing did, but we were disciplined about what we standardised, what we automated and where we invested human engineering time.
• Weeks one and two: discovery and planning.
• Weeks three to six: platform delivery, pipelines and the MVP.
• Weeks seven and eight: hardening, handover, enablement and the roadmap beyond the engagement.
The broader lesson is that speed doesn’t have to come from cutting corners. Sometimes it comes from having fewer corners in the first place. A well-defined architecture, reusable deployment patterns, clear ownership, fast decisions and a platform that removes unnecessary integration points can change how quickly a team moves.
One governed path from source to consumption
Conversations about Databricks often become lists of product capabilities: Lakeflow, Unity Catalog, Databricks SQL, Workflows, Asset Bundles, serverless, AI. These matter, but they weren’t the story. What mattered was what became possible when more of the data lifecycle sat on one governed foundation.
The source landscape included Pro2 and OpenHousing on MSSQL, the Finance database, the TotalMobile DataMart and Azure DMS documents. Ingestion used native Databricks capabilities, including Lakeflow Connect. Data then moved through a Medallion Architecture:
• Bronze retained the raw representation of each source.
• Silver cleaned and conformed the data.
• Gold delivered business-ready data for downstream use.
Databricks SQL supported both transformation and consumption, with Power BI providing the reporting layer. Unity Catalog provided governance, lineage and role-based access across Development, Test and Production. Terraform, Databricks Asset Bundles, Git and CI/CD underpinned delivery.
Because ingestion, transformation, governance, deployment, observability and serving were designed as parts of one operating model rather than separate technology problems, the result was simpler. And a simpler architecture is easier to build, easier to govern and, crucially, easier to own.
Governance from the first pipeline
A question that came up repeatedly around the session was how we built governance in from the start, particularly with sensitive data.
The honest answer is that we didn’t have a choice. If governance is something you plan to add later, you have already made an architectural decision: that it is optional during the period when the platform is changing fastest. We treated governance as part of the definition of production-ready.
Each environment had a distinct purpose. Development was where the team built and iterated. Test was where outputs were validated against agreed criteria. Production was governed, monitored and accepted.
Unity Catalog provided central access control, lineage and auditability. But governance is wider than permissions. We also wanted to be able to answer:
- Who changed this?
- How did that change reach Production?
- Which version is running?
- What happens if a job fails?
- Who owns it?
- How much does it cost, and which project or environment generated that cost?
These are governance questions too, which is why our model extended beyond Unity Catalog into CI/CD, monitoring, tagging, delivery processes and ownership.
SQL as an adoption strategy
One of our biggest architectural decisions had little to do with architecture. We built the transformations in SQL, and only SQL, deliberately.
The client team was SQL-native. They understood SQL, their data and their business. What they didn’t yet have was deep experience operating a modern Databricks Lakehouse. Rather than treat that as a skills gap, we built on what they already knew. The principle was simple: familiar language, new platform.
Technology transformation often asks people to change too much at once: a new platform, language, deployment process, governance model, vocabulary, architecture and set of responsibilities. Then we wonder why adoption is hard. By deliberately reducing cognitive load, SQL became the bridge between the world the team understood and the platform they were learning.
In the early weeks, existing stored-procedure logic was recognised and mapped into Databricks SQL. Then we built together, with our engineers working alongside the client team and knowledge transfer happening continuously rather than being saved for a final handover meeting. By the closing weeks the emphasis had shifted again: recognise, build together, own.
The team also received tailored training, Databricks Academy learning paths, runbooks and access to the repository containing everything that had been built. That is a very different experience from attending a course and returning to the same job the next day.
Upskilling is an architectural decision
This project reinforced a belief we hold more strongly with each engagement: upskilling is not a workstream. It is an architectural decision.
Technology programmes are too often structured in sequence. Specialists build the platform, then document it, then someone creates training, and then the organisation is expected to take ownership. If the people who will operate the platform haven’t meaningfully taken part in building it, we have made ownership unnecessarily difficult. Here, capability building happened inside delivery.
That applied to our own team as well. The delivery model was deliberately small: fractional Solution Architecture providing patterns and guardrails, a Senior Data Engineer responsible for the engineering approach and hardening, and a growing engineer. This was someone certified, with completed learning and internal projects behind them, now taking on their first production-grade delivery.
The wider industry has a problem here. We talk constantly about the data and AI skills shortage, yet job descriptions ask for three, five or seven years of production experience in technologies that haven’t existed in their current form for that long. The answer to “how does anyone get production experience?” is that we have to let them experience production.
That doesn’t mean placing inexperienced people on critical systems and hoping for the best. It means creating environments where people take meaningful responsibility while protected by senior guardrails, automated deployment, testing, governance, code review and well-defined patterns. That is supported responsibility, and it is how experience is built.
Automation shouldn’t remove learning
We didn’t start from a blank repository. We brought repeatable accelerators for Terraform deployment, Unity Catalog structures, ingestion, CI/CD, FinOps tagging and billing dashboards.
Some worry that accelerators reduce engineering. We think the opposite can be true. The question that matters is: what are we automating? If every engineer on every engagement recreates the same Terraform structure, deployment mechanism, tagging convention and baseline monitoring, we aren’t producing better engineers. We are repeatedly spending capacity on solved problems.
Automation should remove the least valuable things to learn repeatedly. That frees engineers to focus on what is specific to the organisation: its data, business rules, quality challenges, security requirements and operating model, and the decisions the architecture must support. Standardisation gave us speed, and it also created safer conditions for people to learn.
One path from laptop to Production
The same philosophy shaped deployment. We didn’t want one process for Development, an informal one for Test and a separate ritual for Production. We wanted one path:
Commit → Pull request → Validate → Deploy to Development → Promote to Test → Verify → Approve → Promote to Production.
Terraform provisioned the platform components, including workspaces, Unity Catalog structures, roles and tagging. Databricks Asset Bundles packaged the SQL pipelines and jobs. Git provided version control, and CI/CD provided repeatability. The same code moved through every environment.
This is good engineering practice, but it has a benefit that is often overlooked: the delivery architecture itself becomes part of the training. An engineer working in this environment isn’t only learning to write SQL. They are learning how production software changes safely, how code review works, how environments differ, how permissions are managed, how failures are monitored and what “done” actually means. Good governance doesn’t have to slow learning down. It can create the guardrails that make learning safer.
“But Databricks is expensive…”
We hear this often, and cost conversations about modern data platforms frequently lack context.
For this workload, the average Databricks running cost has been under £350 per month. That needs a qualification: we are not saying every Databricks platform should cost this. A platform processing huge streaming volumes, supporting hundreds of concurrent users, or training large machine-learning models will have a very different profile. There is no single “Databricks cost.” There is the cost of the architecture and workloads you choose to run on it, and architecture has an enormous influence on that number.
One of the most important FinOps principles is also one of the simplest: don’t pay for something that isn’t doing useful work. In practice, waste accumulates quickly. Development resources stay running. Applications remain online when nobody is using them. Compute is provisioned for peaks that occur occasionally. Jobs run more often than the business needs. Whole datasets are reprocessed when only a fraction has changed. Each decision looks small, until the invoice arrives.
This is why serverless is an important part of the cost conversation. Where it suits the workload, it aligns consumption with actual usage rather than permanently provisioned capacity. But serverless is not a substitute for good engineering. You still need to ask whether a workload should be running at all, switch off unused applications, avoid treating development resources like 24/7 production services, schedule jobs according to real business requirements, and process incrementally. If 1% of your data changed, why repeatedly pay to process the other 99%? Native incremental ingestion patterns were valuable for exactly this reason.
If you only discover cost from the invoice, you don’t have FinOps. Cost optimisation begins in architecture, not in Finance.
From the first pipelines we introduced tagging and billing visibility, which let us ask much better questions. Which environment is consuming the money? Which job or workload? Is the increase expected? Did usage grow because business value grew, or because something changed technically? Can the workload run less often, or on serverless? Is something sitting idle?
Without that visibility, optimisation is reactive: Finance reports that the bill rose last month, and engineering reconstructs why. A modern data platform should let engineers understand the economics of what they build. We already expect engineers to think about reliability, performance, security, data quality and maintainability, so cost belongs on the list. The engineering cycle becomes Build → Observe → Understand → Optimise → Repeat. If we are upskilling engineers to own Databricks, we aren’t finished when they can build a pipeline. They should know how to operate its economics too.
Data for Good doesn’t mean a lower engineering bar
When budgets are tight, or when work is framed as Data for Good, there can be an assumption that ambition and engineering quality are in tension. We don’t believe that. Organisations operating under tighter constraints have less room for technology that is expensive to maintain, poorly governed or dependent on consultants indefinitely.
They need repeatability, observable cost, automation, people who can own what has been built, and an architecture that can evolve without another transformation programme in three years. Doing good with data makes sustainable engineering more important, not less.
Seven principles for repeatable delivery
At the end of her Big Data LDN session, Bianca distilled the engagement into six principles. Having delivered the project, she adds a seventh.
- Say yes to a real scope. Ambition is good, but it must translate into outcomes that can be delivered in the time available.
- Right-size the team. Fractional senior expertise, strong engineers and growing talent can be remarkably effective.
- Start from accelerators. Reuse proven deployment, governance, ingestion, CI/CD and FinOps patterns instead of solving them again.
- Govern from day one. Environments, permissions, lineage, cost attribution and monitoring belong in the first pipeline, not the final sprint.
- Build in the client’s language. Existing skills become the bridge into the new platform.
- Ship through CI/CD. Create one repeatable path from development to production.
- Design for ownership. The real test isn’t whether consultants can operate the platform. Of course they can. It is whether the organisation can operate it confidently when we are no longer in the room.
What success looked like
Eight weeks delivered a production-ready Databricks foundation: three environments, four source systems, priority reporting migrated to the new architecture, governance through Unity Catalog, a Medallion Architecture, native ingestion, SQL-based transformation, Power BI consumption, infrastructure as code, CI/CD, monitoring, FinOps, documentation and training.
Those are tangible outputs, but they aren’t the most interesting measures. The client team could look at a modern data platform and recognise their own SQL skills within it. A growing engineer moved from certification and internal projects into real production delivery. Governance became part of everyday engineering rather than someone else’s later problem. Cost became visible and understandable. And the organisation was left with a foundation that can grow well beyond its first two reporting use cases.
The best platform is the one that lets us leave
As consultants, we have an uncomfortable but important measure of our work: can the client succeed without us? We believe they should be able to.
The goal isn’t to become indispensable by hoarding knowledge. It is to transfer it, establish patterns, automate repeatable work, create guardrails, teach people to operate within them, and then step away.
That is why our biggest takeaway isn’t really about Lakeflow, Unity Catalog, SQL, Terraform, Asset Bundles or serverless, important as they all were. It is this:
The best transformation isn’t the one where consultants leave behind the cleverest architecture. It’s the one where the consultants can leave.
The platform keeps running. The governance is embedded. The cost is understood. The deployment process is repeatable. And the people who remain know exactly what to do next. That is when you have built more than a data platform. You have built capability.
For us, that is what saying yes to Data for Good should mean, and it is a principle we believe belongs on every project.
Want to explore what an eight-week approach could look like for your organisation? Get in touch with the Unifeye team.

