The correlated errors part is what should worry people, more than the error rate. A 1-2% price variance you can price in. Two sources being wrong in the same direction means your cross-check was never a cross-check.
I build on the tax sale side and it's the same shape there. Almost every commercial dataset in that space sits downstream of the same county publication, so when a county posts an amendment or quietly replaces a file, every vendor inherits the identical error at the same moment. You can pull three sources, get three matching answers, and all three are reflecting one bad fetch upstream.
The habit that seems to help is treating provenance as a field. Not just what a record says, but which office it came from and when it was last confirmed. Two records that trace back to the same original aren't two records.
Is there any convention for that in title yet, or is it still per-vendor?
County records are the single source of truth for title in Eastern US States. Title Agents shouldn't and to my knowledge don't, search property records from sources such as the MLS, not for title commitment purposes anyway. For marketing, perhaps, but then the validity of the data is far less important.
However, this is an excellent example as to why title agents need to carefully vet their vendors to ensure where they are getting their data from and to make sure that they are searching county data directly from the county and not a third party who might be sourcing data from other locations.
I am less familiar with the Title Plant model used in Western US, so I can't comment on where they are sourcing their data from or what they are doing to cross reference and validate.
If the county gets something wrong, title may pick up the error, but a properly conducted title search doesn't just rely on the reported data set in the table, they pull copies of the recorded documents theymselves and validate the data based on the document itself.
And sometimes, even the document is wrong. But that's not a new technology problem, that's a problem as old as paper. For example, I've encountered legal descriptions that are wrong. A typo is made once and it's carried on through the record, deed after deed after deed, until one day someone double checks and realizes that it doesn't add up. But there are legal remedies for situations like that. I have seen new technology from Talos Title Ai that will actually make it easier to pick up those sorts of discrepancies in the chain of title, which is a huge win.
Technology is a tool. How you use it is what makes the difference.
That distinction is the useful one, and it's exactly where the tax sale side falls apart.
What you're describing is a discipline: don't trust the index, pull the instrument, validate against the document. That exists in title because the work is per-file, someone carries liability, and the economics support it.
On the tax sale side none of those hold. A buyer is looking at four hundred parcels with a sale date two weeks out and maybe fifty dollars of diligence budget per parcel. Nobody is pulling the recorded chain on all four hundred. So in practice the tax roll becomes the record, and the tax roll is itself a derived dataset. The assessor maintains parcel descriptions on a track that is separate from the recorded instruments, and the two drift. Splits, merges, corrections, annexations. The parcel gets sold under the assessor's description, not the deed's.
Which makes your typo example worse in that context rather than better. In title, a wrong legal description propagates down the deed chain and eventually somebody catches it, because somebody is always reading it. On the tax side there's a second chain running in parallel that nobody reconciles against the first, and the person buying at the sale is the least equipped party in the transaction to notice.
I'm building DeedFlex on that problem so I'm biased about how interesting it is. But the honest version is that I don't think better aggregation solves it. Volume buyers need a way to know which five parcels out of four hundred are worth pulling documents on. That's a triage problem, not a data problem, and pretending it's a data problem is how you end up with three vendors confidently agreeing on the same wrong answer.
Talos is new to me. Is it working off the recorded images directly, or off already-abstracted data?
Talos is working off recorded images directly, but those images are being pulled in by the client from a title search they've done (or ordered from a searcher). But it's the tech driving the document analysis that has real value here, it's easy enough to source the image records to feed into the tech for analysis. I definately think you should reach out and try to have a conversation with them. I'll send you a direct message in Substack to help you connect.
As for the divergence of tax parcel data vs deed chain, I live in Pennsylvania, so I am most aquintanted with our record systems. Here every tax parcel is assigned an ID number. That ID number must be present in every document recorded against the property with the recorder of deeds. While I'm not privy to any information or records used by our tax assessors and tax claim bureaus internally, from the buyer side of things, the only information available to a potential tax sale buyer from either department is basic owner name, address, tax ID number, tax assessment and taxes assessed data. For full property descriptions you have to go to the recorder of deeds offices. So, as far as I'm aware, there is no gradual drift between the two - I wonder if there a big differences between the states where this is going to be more or less of a problem. (But that said, I've never purchased a property from a tax sale, so it's entirely possible that I'm not understanding the full situation here in PA either.)
Honestly, I've always wondered how anyone could buy properties at tax sale without having a full title search done first, but that would be cost prohibitive in bulk.
At any rate, I understand your frustration with the system and totally agree that more or better aggration doesn't solve a problem when the source data is inaccurate.
The correlated errors part is what should worry people, more than the error rate. A 1-2% price variance you can price in. Two sources being wrong in the same direction means your cross-check was never a cross-check.
I build on the tax sale side and it's the same shape there. Almost every commercial dataset in that space sits downstream of the same county publication, so when a county posts an amendment or quietly replaces a file, every vendor inherits the identical error at the same moment. You can pull three sources, get three matching answers, and all three are reflecting one bad fetch upstream.
The habit that seems to help is treating provenance as a field. Not just what a record says, but which office it came from and when it was last confirmed. Two records that trace back to the same original aren't two records.
Is there any convention for that in title yet, or is it still per-vendor?
County records are the single source of truth for title in Eastern US States. Title Agents shouldn't and to my knowledge don't, search property records from sources such as the MLS, not for title commitment purposes anyway. For marketing, perhaps, but then the validity of the data is far less important.
However, this is an excellent example as to why title agents need to carefully vet their vendors to ensure where they are getting their data from and to make sure that they are searching county data directly from the county and not a third party who might be sourcing data from other locations.
I am less familiar with the Title Plant model used in Western US, so I can't comment on where they are sourcing their data from or what they are doing to cross reference and validate.
If the county gets something wrong, title may pick up the error, but a properly conducted title search doesn't just rely on the reported data set in the table, they pull copies of the recorded documents theymselves and validate the data based on the document itself.
And sometimes, even the document is wrong. But that's not a new technology problem, that's a problem as old as paper. For example, I've encountered legal descriptions that are wrong. A typo is made once and it's carried on through the record, deed after deed after deed, until one day someone double checks and realizes that it doesn't add up. But there are legal remedies for situations like that. I have seen new technology from Talos Title Ai that will actually make it easier to pick up those sorts of discrepancies in the chain of title, which is a huge win.
Technology is a tool. How you use it is what makes the difference.
That distinction is the useful one, and it's exactly where the tax sale side falls apart.
What you're describing is a discipline: don't trust the index, pull the instrument, validate against the document. That exists in title because the work is per-file, someone carries liability, and the economics support it.
On the tax sale side none of those hold. A buyer is looking at four hundred parcels with a sale date two weeks out and maybe fifty dollars of diligence budget per parcel. Nobody is pulling the recorded chain on all four hundred. So in practice the tax roll becomes the record, and the tax roll is itself a derived dataset. The assessor maintains parcel descriptions on a track that is separate from the recorded instruments, and the two drift. Splits, merges, corrections, annexations. The parcel gets sold under the assessor's description, not the deed's.
Which makes your typo example worse in that context rather than better. In title, a wrong legal description propagates down the deed chain and eventually somebody catches it, because somebody is always reading it. On the tax side there's a second chain running in parallel that nobody reconciles against the first, and the person buying at the sale is the least equipped party in the transaction to notice.
I'm building DeedFlex on that problem so I'm biased about how interesting it is. But the honest version is that I don't think better aggregation solves it. Volume buyers need a way to know which five parcels out of four hundred are worth pulling documents on. That's a triage problem, not a data problem, and pretending it's a data problem is how you end up with three vendors confidently agreeing on the same wrong answer.
Talos is new to me. Is it working off the recorded images directly, or off already-abstracted data?
Talos is working off recorded images directly, but those images are being pulled in by the client from a title search they've done (or ordered from a searcher). But it's the tech driving the document analysis that has real value here, it's easy enough to source the image records to feed into the tech for analysis. I definately think you should reach out and try to have a conversation with them. I'll send you a direct message in Substack to help you connect.
As for the divergence of tax parcel data vs deed chain, I live in Pennsylvania, so I am most aquintanted with our record systems. Here every tax parcel is assigned an ID number. That ID number must be present in every document recorded against the property with the recorder of deeds. While I'm not privy to any information or records used by our tax assessors and tax claim bureaus internally, from the buyer side of things, the only information available to a potential tax sale buyer from either department is basic owner name, address, tax ID number, tax assessment and taxes assessed data. For full property descriptions you have to go to the recorder of deeds offices. So, as far as I'm aware, there is no gradual drift between the two - I wonder if there a big differences between the states where this is going to be more or less of a problem. (But that said, I've never purchased a property from a tax sale, so it's entirely possible that I'm not understanding the full situation here in PA either.)
Honestly, I've always wondered how anyone could buy properties at tax sale without having a full title search done first, but that would be cost prohibitive in bulk.
At any rate, I understand your frustration with the system and totally agree that more or better aggration doesn't solve a problem when the source data is inaccurate.