First – a bit of backstory: When I formed Tiingo in 2014 I ran into two major issues: (1) data licensing difficulty and (2) data quality issues. Once I purchased the enterprise licenses, the data those vendors sent me was filled with errors.
This post digs into how we approach upstream disagreement, or more broadly, lack of consensus among upstream enterprise data providers. If you think this lack of consensus is small – in July 2026 alone, upstream enterprise providers disagreed with each other on 3,200+ data points before our checks even started. And the problem has been much worse in prior years. This is the heart of how Tiingo approaches data quality and data cleansing.
Here is a chart you will see again – but this chart highlights how often certain enterprise providers can disagree – with periods of market stress (i.e. GameStop 2021) – pushing providers to disagree on up to 2.5% of the values we compared!
Even during calm periods it can be a few basis points per month – but over the years these compound if you don’t handle them with the seriousness the problem deserves.
This is why Tiingo approaches data quality the way it does – by error cleansing and assuming each data point is incorrect.
So this post will dig into the lack of enterprise consensus and what that means for you, how Tiingo handles data quality issues (Composite Indices), and how you (we all) benefit at no additional cost.
How Tiingo handles data quality – an overview:
Tiingo cross-checks every US asset price and fund NAV against 3 to 7 independent feeds, assumes each data point is wrong until it passes validation, and keeps a permanent audit trail of every correction. In practice:
- Upstream providers disagree on roughly 1 in 292 data points. That is 0.342% across 7.9 million equity price and fund NAV comparisons in the twelve months to July 2026, and it ran as high as 1 in 39 in the single most stressed month, February 2021. Every cross-provider disagreement we detect is adjudicated before we publish. A separate class of problem, where several providers carry the same wrong value and therefore never disagree, is what the rest of our validation exists to catch.
- 3 to 7 independent feeds for US assets, so no single provider gets to decide what is true.
- Every data point starts out assumed wrong. It has to pass our checks to be published.
- Composite Indices adjudicate the disagreements, including the harder case where several providers agree on the same wrong value.
- Nothing is ever deleted. Every adjustment is captured with its delta and its timestamp, and nobody at Tiingo can overwrite what was there before.
- 1,000,000+ eyeballs on the data through our customers and white-label partners, all of them able to report anything they find.
- Included on every plan at no additional cost
A disagreement is not automatically an error. But at most one of the disagreeing values can be right, and somebody’s feed is carrying the wrong one. Our job is to work out which, before it reaches you.
Is Enterprise Data More Accurate?
This is one of the hidden challenges practitioners know about in market data. Every major quant fund I’ve worked at or spoken with has a data cleansing arm – and there is little alpha in revealing to your upstream source the errors. It becomes a guarded secret. So many enterprise providers have errors that go undetected because incentive structures lean toward this – also because professionals know there is an error and often have multiple backups, as it’s quicker to build architecture this way rather than wait for upstream resolution.
We can’t make any definitive statements about whether enterprise providers are more accurate as they can vary – but what we can say is that, among our userbase, some of the most powerful reporters of any potential errors are non-professionals. With non-pros, we find they are eager to share their findings, bring a more collaborative attitude, and don’t think of correcting errors internally as an alpha source.
For that reason, our non-professional and enterprise customers alike keep eyes on our data, and given our white label partners, we end up with 1,000,000+ eyeballs on our data, all of them able to report anything they find.
From a data provider side, opening our data can be a reputational risk. Making data accessible to more people means more errors get found and fixed, but it also means more of them get seen. More visible errors do not mean more errors. Every provider in this industry has errors; the ones flowing through Tiingo get found and corrected in the open. And we love that because the goal is better data at all costs, not reputation through obscurity.
When a customer flags a suspect value, the source is almost always an upstream enterprise feed – the same value that provider would have handed you directly, without a second look. We catch the large majority of these before we ever publish. The ones that do get through get found here faster, because more than a million users can report a suspect value.
A small number of upstream values do slip past even our checks – the market keeps producing cases nobody has seen before. When one does, we correct it, we log what happened and who altered it, and the correction sits in the record with its timestamp permanently.
I believe sometimes the knee-jerk reaction is to read a visible error as a sign that we sit lower than an enterprise feed. IMO, what it actually shows is that we audited the point – as this exact error was likely sitting in the enterprise feed before we filter the vast majority of them.
Lack of enterprise consensus – is it a problem?
We ran statistics on disagreements for various periods of history.
The methodology is simply looking at price differences (OHLC) across our providers and the total times we calculated they disagreed. We excluded most OTC here – focusing on NMS securities and funds- to focus on assets that have especially centralized dissemination of data.
A disagreement is flagged conservatively. For a given field on a given security and date, we take the highest and lowest value across our providers, and we flag it only when both of these hold:
(high / low - 1) > 0.001 # more than 0.1% apart
(high - low) > 0.01 # and more than one cent apart
Both conditions have to be true, so rounding differences, sub-penny variation, and small relative gaps on high-priced names never register at all. One comparison is one security-date bar for equities, or one fund-date NAV. The numbers below are a floor rather than a ceiling: anything we count is a genuine, material gap between what two providers say happened on the same day.
Here are the graphs:
Across the twelve months to July 2021, our providers disagreed on 1.161% of the values we compared, about 1 in 86. For the twelve months to July 2026 it was 0.342%, or 1 in 292. The peaks are important here as otherwise market data may fail in the most volatile times – when the window of opportunity is at its widest: in February 2021, during the GameStop period, 2.58% of compared values disagreed. Our calmest recent month, April 2026, ran at 0.228%. The stress showed up on both sides of the dataset: mutual fund NAVs disagreed at 3.2% that month, and equity OHLC bars, which the SIP should keep unified, ran at 1.0% against a 0.40% baseline for the year.
Left unhandled, these disagreements do not disappear – they sit in a vendor’s historical record and make their way into every backtest, every model and every decision built on that feed. This is why fixing them is a personal mission for me.
In absolute terms that is 101,659 disagreements across 8.76 million comparisons in 2020-21, against 27,158 across 7.93 million in 2025-26. February 2021 on its own produced 16,636 of them, more than half of everything we saw across the whole of 2025-26.
The shape of the disagreements changed as well. In 2020-21 a disagreement tended to run across the whole bar: per 10,000 end-of-day comparisons we saw 22.2 on the open, 20.9 on the high, 22.8 on the low and 30.1 on the close. In 2025-26 the high, low and close all fell to roughly 5 per 10,000, while the open rose to 27.5. Roughly two-thirds of the remaining OHLC disagreement now sits in the opening price.
This is the challenge practitioners in the industry face. Even though OHLC of NMS securities should largely be unified in methodology given the SIP, the reality is errors can exist for technical reasons or from mishandled edge cases. While disagreement rates don’t mean error rates, in general, if the methodology is unified, there should only be one value that’s true.
We can clearly see times of stress increase disagreement rates – and when we dug into those stressed periods, the cause was technical errors at the upstream sources rather than genuine ambiguity about the price.
So customers using just one enterprise feed are likely to see higher “deviations from truth” than using multiple. And this is why we use multiple.
Errors – an example of how we identify and fix them
Now that we can see disagreement rates, we then have to dig into them – including erroneous corporate actions. This is where our “Composite Index” process excels at identifying and correcting errors – including when vendors agree on the same error.
When we started running analysis on our adjusted prices in 2014-2016, we noticed massive brief spikes. Digging into it, we found several of our enterprise providers – let’s call them XYZ (this actually affected several providers but for simplicity let’s focus on “XYZ”) – were off by 1-3 days. We put together a spreadsheet with 200+ potential flags – thinking surely this enterprise financial firm would do something about it… but then silence.
And more silence.
2-3 years of silence. I followed up, tickets were re-ack’d and then silence.
Then the ticket autoclosed.
Of course we didn’t wait 2-3 years – but this was somewhat par for the course (XYZ was particularly egregious). We’d report errors upstream to the providers, and it felt like 20% of the time they weren’t fixed on their end, or the fix was unsatisfactory.
How does Tiingo error check corporate actions?
Throughout our product pages you see “Tiingo Composite Indices,” but what are they? This is our approach to handling lack of consensus among financial asset prices in the Equity space. Alternative and OTC assets are a whole other ballgame, which we will save for another blog post – but Tiingo is a major player in OTC and alternative asset pricing (a part of the business we don’t list since it’s more niche and enterprise focused).
We don’t start by looking for better sources. We start by assuming every source is wrong.
With errors like that split, we built our Composite Indices framework. Each data point, price or corporate action goes through a series of checks and validations. Hundreds of millions to billions of data points all being error-checked in real time – going through a series of checks we’ve built over twelve years of being in business and poring over financial data. Some of these calculations are incredibly intense, scaled over a hundred thousand symbols at once.
We then measure this data, see how it varies over time for each upstream data stream, and build self-healing and learning mechanisms that penalize/rank different data streams based on heuristics of change and deep analytics we track.
As an example, let’s bring back the split issue we found above and show how we corrected it.
While we can’t 100% verify the reasons for these errors, the most common source is a misread of a split announcement. Often a split announcement will say it goes “into effect after the close on the 17th” – which really means the ex-day is the next trading day, the 18th in this example. We think corporate action providers misread this and mark the ex-day a day early.
Regardless, because we integrate with 3 to 7 independent feeds for US assets, we can catch a misidentified split, and even have checks for when multiple providers get it wrong. Our system then has a methodology (refined over a decade) to determine which corporate action is correct – including when there is consensus around the wrong point. In the XYZ case we corrected the affected history on our side, logged every adjustment with its timestamp, and added a check so the same misread would be caught automatically next time. I’ll leave the specifics of how we establish the true ex-day out of this post, since that machinery is a large part of what we’ve built.
Auditability of Data is Everything
What we described above is one check. In production we run many scans and tests catered to each data point. This computationally intensive process is why we don’t make data available the moment we receive it in raw form. We process it, analyze it, and then release it.
After the tests are run, we can pull statistics each day on how each data point performed, errors passed, errors corrected, and which data points ran clean. No data point is deleted, but updates are tracked so we can see how quickly data was corrected and improved.
And in the rare cases where there was no consensus, or all providers were wrong, we have a human override in which a team member comments on what happened, what the error was, and what data points we manually adjusted. This is so we can keep track of human performance as well.
Everything we do is measured, every error is measured, and in some cases we run aggregate statistics and make procurement decisions based on who is underperforming.
But sometimes a provider that underperforms overall is strong in one specific area, and we keep the relationship anyway – the biggest thing is predictability of errors, because predictable errors we can correct.
All of this comes back to the mission. If solving inequity and “Actively Do Good” are paramount, we have to be clear-eyed that errors in upstream market data are inevitable, while accepting them or concealing them is not. What matters is how systematically they get detected, corrected and documented. There is no error-free financial data anywhere in this industry, and once you accept that, the question then becomes – how do we minimize the time and friction from a report to a resolution – and the only way that happens is if we open as much as we can to all.
Going forward, we are going to make how many disagreements we handled public, including human interventions. And when customers have helped us identify them, we will give them shoutouts (with permission) for helping make the dataset better for all.
With this I close: by being open, I believe our data will continue to exceed enterprise solutions – which due to customer composition, and gateways, often remain more closed. My goal is to encourage openness, and to repeat, more visible errors do not mean more errors. My argument is the opposite: a company that publicly shows how many disagreements are handled via upstream enterprises, and how many errors it corrected, is more transparent and more reliable.