Skip to main content

2025 marks 20 years since Lynchpin started trading, and 25 years for me personally working in the field of data and analytics (isn’t it nice when the numbers line up so beautifully!).

There’s been plenty of exciting developments and change over those decades, but also some unbending truths that haven’t really shifted an inch. In the context of the latest AI wave, it’s well timed to reflect on some of the key shifts and anchors over the past couple of decades, and how they might project forward based on that experience.


Web Analytics Goes Full Circle

In 2000, web analytics was all about counting – vanity metrics were popular with Venture Capitalists – and the ability to track end-to-end marketing performance was rarely even articulated as a need.

The gold standard of “how many visitors?” was key to pumping up valuations, but of course nobody could agree how “a visitor” should be defined. The Audit Bureau of Circulation – who were more used to auditing newspaper circulation figures – stepped in to provide an industry standard that could be calculated in 3 completely different ways, and my first foray into web analytics was writing Perl scripts that would churn through web server logs at an early dot com business to give the magic audited number that would keep the investors happy.

In the early noughties, the web analytics industry was exploding with platforms that put not very useful and not very performant front ends on top of clickstream data. And then that market consolidated in what felt like a heartbeat when Google released a free analytics tool in 2006. But to answer any real business questions, the only real solution was still to stick the raw data into a database and start to build your own data model.

At Lynchpin we bit the bullet early and literally racked up database servers in a datacentre (this was pre-cloud) to answer those key questions: what was the real customer journey across display and search (in 2008), how did publishing recency and frequency affect search visibility and hence reach for newspapers globally (in 2009), how could you model the long term impact of content on long term customer acquisition (in 2014).

We also learned early on that to effect actual change you needed to give people tools, not just reports, that they could use in their day-to-day decision making. We were super early adopters of Tableau back in 2008 (Power BI didn’t appear until 2015) to arm our clients with role-focused tools that gave them data-driven recommendations and scenario planning aligned with the levers they could pull around trading, merchandising, pricing and budget allocation.

Fast forward to today, and it does make me smile a little to see the de rigueur model for enterprise digital analytics is converging on just sticking the raw clickstream data into a data warehouse and taking it from there – accelerated not least by Google Analytics 4 being arguably a step back in terms of user interface coupled with a convenient free stream into BigQuery.

And with the shifts in consent for digital measurement and the stealthy introduction of modelled data into the platforms, the old question of “what exactly is a visitor” still feels very relevant.

So, a complete full circle? Not quite. What has changed is the accessibility and affordability of compute (which is a recurring theme throughout this reflection): a small cloud processing bill rather than a rack of servers as the entry point for really getting stuck into digital data. What hasn’t changed is the importance of getting the architecture and data model right and the challenge of successfully aligning the outputs with the business outcomes.

Measuring Behavioural Shifts

It’s very human to assume the rest of the planet thinks and behaves as we do, and a consistent danger point in analysis has always been assuming the average represents everyone and then missing the more subtle and slowly shifting tides of different underlying segments.

Throughout the twenty-tens there were at least 5 “This is The Year of Mobile!” pronouncements I can recall; in reality, none of those years were “the year of mobile” and every year was “a year of slightly more mobile”. The percentage shares slowly shifted, at different rates in different demographics, and at the same time the internet started splitting into Apple/Safari/iOS “privacy first” versus Google/Android/Chrome “advertising first” worlds.

I suspect a similar thing is about to happen with shifts from keyword search to prompts to agentic delegation, set against potential cravings for more direct engagement with “authentic” humans. I’d certainly be surprised if it’s a uniform transition across all audiences in all contexts.

Just like the “year(s) of mobile”, the businesses that measure and understand the nuances of these shifts in terms of their customers will be the ones that will successfully tap into the opportunities and manage the threats. And each new mode of engagement drives an important challenge to address in terms of how to measure that engagement in the first place.

Follow The Money

Or at least watch it. Flows of capital determine rates of progress but also define (sometimes inflated) expectations of value that need to be realised to give a return on that capital.

When Uber was doing its initial market entry, a friend used to cheerily remind me that for each unfathomably cheap ride a generous VC had basically put their hands directly into their pockets to pay a chunk. It was all subsidised to grab the market… until it wasn’t anymore.

Similarly, every new wave of technology progress is heavily subsidised… until it isn’t anymore and the investors that won the market share race want the return for their risk: cloud compute, data lakehouses, customer data platforms, large language models.

A comment that stuck with me from an early debate about the relative merits of the different cloud providers was the observation that all of them were ultimately just reselling electricity packaged up into different units of consumption.

That comment feels even more relevant now that the cloud providers are literally building their own power stations to fuel AI usage, and the question of the sustainability of our accelerating compute addiction is an increasingly real environmental issue. Quantum computing is genuinely exciting because of its potential to rewrite the compute vs electricity relationship, but given how long term and uncertain the pay-off, will the capital flows deliver it in time?

Regulatory Lag

In contrast to the capital flows that seek to define the future, regulation has often struggled to shut doors that have been blowing wide open based on the rate of change and legislative lag. Reflecting on the past 20 years it’s often felt like we’ve been doubling down on trying to fix the last but one issue in data privacy while newer and much bigger risks run unchecked.

To start on a more positive note, I’d venture GDPR has stood the test of time as actually being a decent piece of legislation. Mainly because the principles of GDPR all align well with good principles of data strategy – define your use cases and process data with clear purpose.

By contrast, the legislative failure to properly integrate e-Privacy and GDPR in Europe has been a slow-motion car crash of incomprehensible cookie consent dialogs. I don’t think when Tim Berners Lee invented the world wide web his vision was that everyone would be reviewing and accepting a new set of legal terms and conditions for every hyperlink they followed. The reality is that cookie consent should have been a browser standard and the industry failing to pull together on that or put aside some of their vested interests destroyed the user experience.

Meanwhile, other potentially far more intrusive data collection (mobile apps and location spring to mind) ran relatively unchecked, and we ended up in the weird situation where Apple became the self-appointed global regulator of consumer privacy on their own devices.

I can’t see AI regulation being anything other than a horse and stable door type situation in the current geopolitical climate. Which arguably pushes the responsibility back on us as an industry to wrestle with the challenges and lead more effectively in terms of protecting consumer data and business knowledge.

Waves of AI

Around 2018, analytics had a big AI Wave centred around machine learning, and suddenly every analytics and marketing platform was “AI Enabled”. What was new? Nothing much in fact: most of the algorithms being deployed (e.g. to do some basic segmentation or forecasting) had been around for decades, and some were simply rebranded centuries old maths and nothing to do with an AI revolution at all.

The real AI wave around machine learning eased in more subtly from that point, fuelled by the availability of compute, and specifically the proliferation of cloud and open source. As soon as well-maintained machine learning toolkits were openly available as Python packages that could be run in the cloud at negligible cost, the previous platform barriers to entry (expensive servers and licenses) started to evaporate.

With those platform barriers to entry swept away, the increasing gap was practical skills and experience –moving models beyond proof of concept and into production and integrating with the business process. Interestingly, many of the machine learning models we successfully deploy for aspects such as pricing stop getting called AI by the client as soon as they are operational: when it works it’s just “automation” or “optimisation”.

So, does that previous wave of AI help to anticipate the impact of the current wave around generative AI on data and analytics? Certainly the availability of compute is again spearheading progress, and aspects of open source (or at least open model weights) are fuelling the competitive bleeding edge.

Interfaces are Key

What is different this time is the accessibility of the human interface – in theory the final technical barrier to entry swept away. In a world where anyone can prompt to an apparently unrestricted range of outcomes, where does data, knowledge and experience fit in?

Analysts need accessible, timely and well-structured data that they can trust to base analysis on, and models and agents are no different. And lots of businesses still struggle to arm their human analysts with the right data at the right level of detail to do their jobs effectively.

RAG (Retrieval Augmented Generation) allows LLMs to tap into live first party data (e.g. live transactions) but can only be as good as that first party data in terms of quality and availability.

There are also opportunities to turn RAG on its head and use LLM research to feed closed loop machine learning models with additional research context (e.g. new features for forecasting), if there is a good approach to ongoing testing and optimisation.

Finally, emerging (albeit competing) standards like Model Control Protocols (MCPs) give broader interfaces for LLMs interacting with other systems, but in an enterprise need very well-defined governance to manage risks and avoid data breaches.

These behind-the-scenes interfaces are going to be the critical success factors for enterprise AI implementations. I believe this is the biggest built-in hedge in data and analytics right now: humans and machines need to be fed with well-structured and trusted data, and the data strategy fundamentals are very similar irrespective of the consumer.

This latest AI wave just means the businesses that aren’t on top of their data will fall behind much more quickly.

Context is Value

The most important and underrated analyst skill has always been the ability to interpret data in the context of a business and all its nuances: the real bridge between insight and execution and outcomes. But context requires the bringing together of knowledge, and interpretation requires critical thinking.

Can GenAI bridge, or help to bridge, that context and critical thinking gap? Critically, models can only feed off the data made available to them, and when every vendor is sticking a GenAI layer (and cost) onto their platform there is a real risk that valuable proprietary context gets siloed within those platforms.

And because capital flows determine that only a very limited number of players can win the race to train the best overall model, ultimately a limited number of models will end up licensed to other vendors to repackage, further homogenising those layers of siloed context.

It’s AI bloat, and eventually businesses will need to consider the multiplying cost impact. Or, more importantly, unless businesses own and consolidate their own proprietary context that leads to executing on those platforms with a competitive edge (based on better data and hence knowledge), that bloat is simply a race to the bottom in terms of performance rather than an investment.

So, while the current AI wave might be focused on “AI everywhere”, the tenants of data strategy and what makes proprietary insight valuable are not going anywhere. Context is value, and businesses will ultimately need to train models on their own consolidated data to be competitive and prosper.

The Next 20 Years

I wonder how many AI waves there will be over the next 20 years. Will we notice them as they pass, or will it be the more subtle shifts underneath those waves that turn out to be more impactful when we look back?

For those of us that had the fun of working through the dot com boom and bust of the early noughties, it’s hard not to see at least some parallels with the investment patterns and short-term expectations right now. But it’s also an important reminder of how the impact of a technology revolution can sweep in beneath the initial hype once that hype has subsided and turn entire industries on their heads.

Ultimately, interfaces are key, and context is value in data and analytics. I believe this will continue to define the relationship between smart humans, smart technology and real business outcomes in this industry.


Privacy Preference Center