Justin Heller (Opener):
I’m going to be a little controversial. I don’t necessarily think the gap is data when it comes to AI. Someone asks a question about the current strategy and gets an answer that’s partially informed by things that were 25 years ago. Sure, that might be correct, so it may be accurate data, but it’s no longer relevant data. You can evaluate the quality of a fact because you can test it against rules or a rubric, beliefs are different. Don’t sponsor a data initiative as a data initiative.
Saket Saurabh:
Hi everyone. Thank you for listening to another episode of Data Innovators and Builders. Today I’m speaking with Justin Heller. Justin is the former SVP and Chief Data Officer at Synchrony Financial. Justin, thanks for chatting with me today.
Justin Heller:
Thanks for having me.
Saket Saurabh:
Okay, Justin, so let’s start with a little bit of background about yourself.
Justin Heller:
Sure thing. As you mentioned, I’m the former Chief Data Officer at Synchrony. I was Synchrony’s Chief Data Officer for nearly 11 years. I was brought in after Synchrony had separated from GE Capital Corp. For those not familiar with Synchrony, it’s the largest provider of private label credit cards. I had the pleasure of working there, both establishing the data management program and bringing it through to a fairly high level of data management maturity, to protect the organization by mitigating risks, and also to help enable and drive growth.
Saket Saurabh:
Let’s get into the data management practices. Give us some context, in your experience, on the foundational things that help chief data officers or data teams deliver good outcomes over time, especially given that evolution you’ve seen from analytics all the way to AI.
Justin Heller:
Yeah, I think first of all, the mindset is really critical. Data management is not a project, it’s a set of processes, and these processes have to exist to enable other business processes to succeed. One of the first things that’s really critical early on is that you’re not establishing a program with the goal of bringing in tools, and you’re not establishing a police state that prevents people from doing certain things. At the end of the day, data management is a set of processes that help instill trust in data. It’s meant to define who has accountability, not at a system layer, but based on the type of data it is.
Saket Saurabh:
Right.
Justin Heller:
This kind of gets back to the role of applications. Technology controls prevent bad data from being captured in the software that’s being created, a drop-down field is really important. Data quality’s role is to detect if anomalies have happened, so you need a detective set of capabilities to identify where controls are missing or failing, because ideally, if you properly prevent bad data, there’s nothing to detect. If you’re not allowing bad data to come in, you don’t have a data quality program, data quality measures the defects.
Saket Saurabh:
One of the top-of-mind things, as you’re talking about enablement, is AI use cases. Maybe a year and a half back, everyone was saying AI would change so much of enterprise operations, and then people started hitting data challenges in making that possible. Give us a picture of how these data practices apply to AI, or where the gap still needs to be closed to deliver enterprise AI.
Justin Heller:
Yeah, I’ll be a little controversial. I don’t necessarily think the gap is data when it comes to AI. I think the gap is moving from a use-case mindset to a value-creation mindset.
Saket Saurabh:
Yeah, I’d very much agree, and I’ve seen that in practice. For example, in use cases involving leasing documents that have information in them, and similar information in a database from the leases. When someone asks a question, both are data sources, so which one do you trust? And in the case of unstructured data, people use citations to say, “I’ve given you the answer from this document,” and that document might be 10 years old and no longer relevant.
Justin Heller:
Or, worse, it’s a copy, not the authoritative version. We talk about systems of record and relational systems, but we generally don’t have duplicate copies in those relational systems. We might have data that should have been purged because it’s in a temporary structure, but very rarely the same record duplicated, referential integrity prevents it. But in an unstructured world, you can have multiple copies of the same thing, and it’s hard to declare which is the true record, the system of record. That’s why I think records management is going to see a resurgence, shifting scope from just retention to proper, simplified categorization, labeling, and disposal.
Saket Saurabh:
Yeah, absolutely. One thing you mentioned earlier, Justin, was metadata and its different types, and its importance. I think the benefits of metadata have started to show in AI as well. I’ve also noticed that documents in an enterprise are often another source of metadata that can add more context to what a certain attribute or data point means. What’s been your experience or thought process on how this metadata field is growing and bringing relevance?
Justin Heller:
Yeah, much like relational or structured data, you have business metadata and technical metadata, and it’s the same with files. The technical metadata would be file name, creation date, location, owner, last update date, or last access date. This is where operational metadata becomes really important, similar to the information security discipline of monitoring who has access to what and when something was last accessed, which matters for determining relevancy. Just like in structured data, you want to know which structures are most queried.
Saket Saurabh:
Right.
Justin Heller:
So I think you’re right, it is similar in that you’re dealing with both technical and operational metadata. But business metadata in the unstructured world becomes much more challenging, because we don’t always understand how to develop context for a file, since files contain multiple types of information. In structured data it’s simpler, a data element, like customer name, is one type of data. But a file might contain multiple types of data, which is why it’s important to think about what process created it, that’s additional metadata, and why it was created, that context, so you can categorize it. Is this a tax record? Is it used to run a business process? Is it a compliance record?
Saket Saurabh:
Talking about relevancy, one thing much discussed right now is the context layer. Unlike a human using data, where you have context and can say “this document is old,” or discern between two different data sources and know which is the right one to use, an AI agent doesn’t have that knowledge, and the context layer becomes the source of that information: what the data means, what relevance it has, what quality it has, and so on. What’s your take on companies starting to build these context layers? Do you think that’s just an evolution of metadata, or something to be thought about fresh and differently?
Justin Heller:
No, I think it’s an evolution of metadata. We talk about semantic layers, but at the end of the day it’s still data, data about data. What’s important is how that metadata is curated. Where does it come from? Is it self-documenting, discovered automatically, or does it require enrichment from people? Business metadata requires people, because any term could mean something different in different organizations. It’s that combination of people adding contextual relevancy, captured alongside the other types of metadata, technical and operational. This is where human in the loop becomes really important. Machines can’t necessarily uniquely identify context, the context is framed by the manner in which something was created or used.
Saket Saurabh:
Yeah, as you said earlier, organizations that build solid foundations for operational and analytical use cases are already benefiting in AI applications. I think what’s added to the mix is thinking about how to manage unstructured data in a systematic fashion as well.
Justin Heller:
And it’s all about, once you create that foundation, not creating it for the entire enterprise, but to solve a specific business problem first. Then you have an MVP, a minimum viable product, which is really important. The goal for the next problem isn’t to build another bespoke solution, but to expand the current one, enrich it, and onboard the next process. If the mindset is that you have to build something new each time, you’ll never create a foundation. But if you build it, then onboard other platforms and capabilities onto it, and continuously improve and enrich it, you’ll incrementally get to the foundation. Not every business process is ripe for AI, so always think of it in those terms.
Saket Saurabh:
So, Justin, you mentioned MVP, and one of the things top of mind is building these AI MVPs quickly, experimenting fast. I wanted to ask, when we think about data governance, quality, and controls, alongside this push to move quickly and build use cases, how do you see those competing with each other from a data perspective?
Justin Heller:
Don’t stand in the way of innovation. If you’re experimenting, let it happen, do a risk assessment. This is where human in the loop becomes really important. I think most organizations are in this exploration, experimentation, pilot phase, that’s fine.
Saket Saurabh:
Right.
Justin Heller:
Once the pilots are done, before they move into production, there should be an assessment that includes data. Are we happy with the outcomes? Do we trust the data? What risks surfaced, what problems arose, and then address the problems that arise. If the problems don’t arise, then don’t stand in the way of the business. But assessment is critical, that’s number one. If an organization isn’t assessing before rolling into production, they’re absorbing risks that may eventually be greater than their risk appetite. I think the challenge isn’t the pilots or the experimentation, it’s the ability for these things to be deployed enterprise-wide, where an organization can answer what value they’re getting from their AI investments. So I don’t think it’s a data governance issue, I think it’s an AI value-creation issue, and there’s a separate operating model needed.
Saket Saurabh:
Okay, I think that’s great advice for those building AI use cases who of course want to keep the governance, but it applies to when you go to production, wanting sandboxes and controls, or the simplicity to at least try and move innovation forward without getting in the way, as you said.
Justin Heller:
Yes. Governance should not impede progress. Governance should help protect things from going off the rails.
Saket Saurabh:
Justin, this has been a great conversation, thank you so much for taking the time. You’ve been an exceptional CDO, one that’s had a long tenure, and you also owned data engineering as part of that, right? That’s not…
Justin Heller:
Well, actually, no, I did not own data engineering.
Saket Saurabh:
Okay, oh, I thought you owned engineering.
Justin Heller:
No, I didn’t own it. I had accountability for it, but I didn’t directly manage it. As a CDO, you have accountability for many things you may not directly manage.
Saket Saurabh:
Any words of advice for data leaders who want to become CDOs, or who are in that seat today?
Justin Heller:
Yeah, I’d say don’t sponsor a data initiative as a data initiative. Seek to help other initiatives remove impediments through data management capabilities. Data is not a project.
Saket Saurabh:
Right.
Justin Heller:
And I think those with a project mindset, versus a capabilities mindset, may struggle more than those who look at it as, “I’m here to help other people achieve their goals.”
Saket Saurabh:
Yeah, absolutely, and that’s a great way to wrap up, because data is indeed the means to the end. We’re all here to make business outcomes happen and be successful. It’s an important role to play as enterprises try to become AI-first as well. Justin, thank you so much for taking the time today, I really appreciate you sharing your insight.
Justin Heller:
My pleasure. Thanks very much for having me.