Document Graph Markup Language or DGML is an open source document standard.
It's designed to turn complex, unstructured business documents into verifiable machine readable data for AI systems.
Docugami and Nvidiagram have partnered to open source DGML, transforming messy business files into trusted verifiable data for AI agents.
And while Docuga generates the DGML standard from unstructured documents without requiring templates or heavy model training.
Nvidia brings the on-chain verification.
Now, joining me to discuss this and more is Gene Pauli, co-creator of XML and CEO of Doccuami, and Patrick O'Mara, the chairman and CEO of NVNAM.
Xi, uh, good morning to you gentlemen.
How are you guys today?
Good to be with you.
Well, Jean, let's, uh, Jean, let's start with you.
So what is DGML and why does it matter?
Right.
Well, today everybody wants to provide documents to AI agents, and the question is how, and you need two reasons basically.
Number one, it has to be economical.
And so today it's actually not when you try to give PDF files and all of this becomes very expensive and DGML really exposes the higher ratio of extraction accuracy.
Per token, and the second one is trust, because at the end of the day you need to be able to trust and to multi-value parties.
Uh, the, the transactions and vidium chain and blockchain is the best place where you can hash DGML fragments to enable this bidirectional trust.
And John, this question for you, why this partnership right now?
Well, Elvinium and Pat are really at the forefront of the of the overall finance and putting real assets into the market and so my friend Pat and I, we have been talking about this since 4 to 6 months.
And Pat will explain to you by himself that many, that, you know, Invinium has been in this business since many years in a very, very successful manner, but Invinium needed the standards such as EGML to really expand the market in amazing ways, Pat.
Yeah, sure, that, it, it, it, this is a win-win for both of us, right?
We are, Invenium is uh an unstructured document, uh, management solution, right?
Where we help people manage all of their unstructured documents, uh, as a, in, in a decentralized manner and relate those to, uh, data-rich, low-frequency trading assets, so, private market assets.
And so people can understand, they can keep their asset ready for sale 24.
7, and we are that operations management for this unstructured type of documentation, providing proof of origin process and state of data as automation is here.
Now, as we do this, we needed to go beyond just to prove this was the document, but we're in that document.
And so the ability to not only index and vector embedded, but also semantically chunk it.
And have those semantic chunks anchored on chain and be provable, we turn a business document or a contract.
Into a registry that multiple people can trade to without leaking any data.
And this is very key, this, this data sovereignty.
And so, we now have uh, uh, an enterprise solution, an open-source solution in the form of DGML and then avenium chain below it.
And so we have the full architecture for private market assets to trade with full attestation, and the agent that interacts with the data can prove origin, process, and state of data.
So Patrick, this question for you.
Tell us why on-chain verification is useful at this moment.
Well, so right now, you know, you need to break open the black box of AI, right?
Um, if an agent is gonna go and, and use data, uh, and, and it needs to do so at scale and at speed, you need to have a mechanism for it to prove, I got the right data from the right location, um, and I can, I can understand not only the origin of that data, but the process, meaning who improved it, and the state, meaning it's the most recent version.
And using DEGML that agent can inter, interface or, or interact uh with, uh, uh, uh, an unstructured document, a lease, a contract, uh, a, a, a, a subscription agreement, a limited partnership agreement, and say, I got this data from this spot, and that is the element needed for us to process this data and either make it.
Distribution, a capital call, allow something to be traded, and we're giving that attribution not only to the document, but we're to the actual piece of the document and where that lives.
And because this is open-source, we want this to be the native language of agents, right?
Uh, Jean Paulli is the primary, one of the three authors of XML, the primary language of the internet.
And, you know, the, the, the main contributor to that.
And what we have now is uh a new document graph markup language that we think will be the primary language of agents as they interact with unstructured data, which is where we live.
And this is the lingua franca of private markets.
If you wanna buy, sell, private market assets, you need to understand the state of that asset so that you can understand the value of it, and then you can run the waterfall, that carry calculation, that expense allocation, and then. that capital stack daily.
That agentic function needs attestation, auditability, and transparency.
And that's what the on-chain process does.
And DGML is gonna be the language by which the, the agent interacts with all of those unstructured documents, knows the right one, and then literally can deliver it to an end third party, to another agent, or to an end system.
And this is kind of the whole uh tech stack bundled together.
So you know, you asked why exactly now because now it's where we are done from going from AI as an experimentation to have actually agents really doing real things at scale.
In order to do real things at scale, you need to be able to have business documents, the one that Pat was just describing, literally helping.
And how do you, are you going to do that at scale without having AI agents really burning so much tokens?
I mean, you know, token economy is.
Real dollars, right?
Like, and so all of a sudden having it, uh, you know, being able to be using a very good format that we, we believe could become a standard across the, across the world, and we are, that's why we are actually asking, asking many people around the world to come to to the GitHub repository and the place we are doing the standardization to help uh with us test those financial uh scenarios that Pat was describing.
And Patrick, this question is also for you.
So take us to the real world use cases and what will this mean for things like valuation, asset managers, and AI agents working with hard to value private market assets.
So, I, I, I'll start at the end of the last question, and I'll work backwards to what you, uh, uh, asked, which is, if, if we think about, uh, agents and their need to deliver something which is audible, particularly in regulated industries, you have to be able to recreate, uh, that calculation and, and not just have something that's Probabilistic.
You need to have something that's deterministic, where you can say, I understand where it got the data, it got the right data from the right location.
It was improved by the right third parties, and then literally it delivered an outcome.
This is what we're delivering, uh, so that we have the ability for that agent to prove origin, process, and state of data.
That's the Invenia Mayo solution.
But as it interacts with that data, it's gonna say, I know that I'm working on the right document, but within the document, to the degree that it's semantically chunked using DGML, suddenly that document, it becomes a registry.
And that registry, uh, can be traded on by many, many different third parties, all relying on the one clause that it says, this is the piece, I'm trading on this based on the elements of, uh, the, all of the documentation around this asset.
And what we're gonna do is we're gonna get to the point where we see interval funds with daily subscriptions and weekly redemptions.
We're not gonna see quarterly, 2%, 3% gates.
We're gonna see much greater flows.
And that begins primarily with better data, right?
And as we can shrink, right now, the average discount for real estate is like 30% if you're selling ahead of time, if you're an LLP trying to exit.
The primary reason is a lack of marketable And market information on those assets.
And what we're doing is we're giving you the ability to surveil that asset on a daily basis, agentically, understand it, rerun that, again, that waterfall, carry calculation, expense allocation, and then reprice the capital stack on every asset in a portfolio, so you can trade a portfolio of multi-family housing, uh, light industrial warehouses, uh, commercial real estate, and what.
All of a sudden, you start to do is see a change in these stores of value, where we're gonna see public and private market ETFs or ETPs where we're, we exchange-traded products or funds that are being traded, that are a collection of these like assets that are being repriced every day.
And if they're repriced every day by these agents, you need to be able to audit those agents and the work that they're doing.
And not only do you need to audit where they got it from, it's the right data for the right building.
You need to be able to go and say, the canonical documents that govern these, I'm trading based on this, and DGML is with the semantic uh chunking with each chunk anchored, not where the chunk is on chain, but that zero-knowledge proof, and Suddenly, we have the ability to trade on the, the constitutive elements of, or the governing elements of those uh portfolios.
This is enormously valued.
But in addition, suddenly we're gonna be able to use AI across entire portfolios and do large dataset pattern recognition and correlation without leaking any data.
This is important.
And that's kind of where what's important with EGML is, as you saw with Pat, he's describing a lot of different document types.
And so when you talk about these kind of transactions at scale across, you know, entire portfolios, you have rentals and liabilities and loans and mortgages and insurance claims and all of these things.
Those are very different document types.
And now that's a lot.
And then each document type you can have hundreds and thousands of millions of pages, and you need to be able to equally understand semantically.
And this is the most importantly GML.
You have what we call semantic context, meaning smaller pieces instead of having hashing entire documents, you hash small.
Pieces in the blockchains, as, as Pat says, we don't put the data on the chain, we put just the, you know, the, the, the, the, the, the signature and then this way you can go back and go later, you go back and later you go inside of the documents and you just double check that all was like you left it 3 months ago or 2 days ago.
Well, gentlemen, thank you so much for joining us today.
Definitely interesting to see how this partnership is going and where, where it will go.
Thank you guys for joining us.
Thank you.
Thank you.
Have a good day.
Thank you.