---
title: "AI Agents Need Clean Data | Cohesity CMO and Kaggle Founder, Sumble CEO"
episode_number: 6
canonical_url: https://enterprisealignedai.com/episodes/marketing-ai-with-cohesity-and-sumble
md_url: https://enterprisealignedai.com/episodes/marketing-ai-with-cohesity-and-sumble.md
last_updated: 2026-08-19
duration: "43 min"
youtube_url: https://www.youtube.com/watch?v=JVO3V2rORhg
spotify_url: https://open.spotify.com/episode/2Wu9brspVpOyo9K4v0dt2g
apple_url: https://podcasts.apple.com/us/podcast/ai-agents-depend-on-clean-data-cohesity-cmo-kaggle-founder/id1890781384?i=1000784472053
guests: ["Carol Carpenter", "Anthony Goldbloom"]
tags: ["Marketing AI", "Go-to-Market", "Knowledge Graphs", "ROI & Measurement", "Data Privacy", "Organization and Talent"]
---

# Ep. 6: AI Agents Need Clean Data | Cohesity CMO and Kaggle Founder, Sumble CEO

**Summary:** Host Aparna Sinha sits down with Carol Carpenter, Chief Marketing Officer of Cohesity, and Anthony Goldbloom, CEO and co-founder of Sumble, on the function that adopted AI first and is furthest along with it. Carol runs a marketing org that is a heavy user of enterprise LLMs, a builder of custom agents for translation and brand compliance, and a demanding buyer of third-party AI tools. She gives the numbers behind each: a 60% target reduction in translation cost, lead follow-up moved from 20% to 80% of MQLs, and top-of-funnel gains in the 10 to 20% range. She is equally direct about where AI has not delivered, and about how much human design work sits behind every agent that works. Anthony sold Kaggle to Google in 2017 and now builds the grounded data layer that go-to-market agents run on. He explains why a language model cannot tell you which team inside an account is ready to buy, what a knowledge graph adds at the team level, and why the work is mostly cleaning data. They disagree productively on whether an enterprise should build its own knowledge graph, and Carol is unmovable on one point: her proprietary data does not leave the company.
## Why listen

Marketing was first to adopt AI and is now far enough in to have real numbers rather than pilots. Carol gives hers with the caveats attached, including where the gains are only 10 to 20% and where the setup work was harder than she expected. Anthony gives the clearest available account of why grounding matters: language models handle the reasoning, but the facts about what companies are doing have to come from somewhere. If you are deciding what to buy, what to build, and what data to hand a vendor, both sides of that call are argued here by people who have made it.

## Key takeaways

- Enterprise B2B took 57 touch points to close a decade ago and now takes 152, which is what makes orchestration a data problem rather than a volume problem
- Lead follow-up moved from 20% to 80% of MQLs through a combination of a few more SDRs and better tooling, not headcount alone
- The AI translation tool is targeting a 60% cut in translation cost, built on OpenAI with brand guidelines, tone, cultural nuance, and a memory of prior translations
- Image generation and the automated customer journey are the two places AI is not yet a game changer in marketing
- Agents do not automate the customer journey on their own; a person has to design the sequences and set the baseline before an agent can learn from it
- A language model is good for the thinking, but the thinking has to be applied to a hard-grounded corpus that knows what the world's companies are doing
- Sumble maps tech stacks and buying windows to individual teams inside an account rather than to the company, which is the entity CRMs are missing
- The moat is curation and depth: 70 million companies in cold storage, 3 million active, and roughly 40,000 that buyers ever look up, where the data is kept pristine
- Building the product is 80% cleaning data, including disambiguating job titles that drift from SalesOps to RevOps to GTM engineering
- Putting all go-to-market data into a warehouse gets roughly 70% of the value of a knowledge graph, because language models are good at SQL
- Proprietary first-party data does not leave the enterprise, even if sharing it would make a vendor's agents smarter
- Hiring for AI means asking about the candidate's latest prompt and the problem behind it, not whether they can write one

## Timestamps

- 0:00: 152 Touch Points to Close One Enterprise Deal
- 1:04: Meet the Guests: Cohesity's CMO and Sumble's CEO
- 2:49: How Cohesity Uses AI
- 4:09: Founding Kaggle and Sumble
- 5:50: Where CMOs See Real ROI from AI
- 7:33: Busting the Hype of AI in Marketing
- 9:04: Agents Require Humans: Don't Underestimate the Work
- 11:23: Cutting Translation Cost 60% and Moving MQLs 20% to 80%
- 14:55: Sumble Provides Team-Level Tech Stacks and Buying Windows
- 17:43: Sumble's Moat: Pristine Data on 40,000 Companies
- 21:57: Data Privacy Is Paramount for Enterprises
- 24:06: Building Sumble Is 80% Cleaning Data
- 27:02: Build vs Buy: Should Cohesity Build Its Own Knowledge Graph
- 31:22: Cohesity's Brandi Agent and the Pre-Event ROI Calculator
- 36:51: Organization and Talent: Defining the GTM Engineer
- 38:37: Hiring for AI: "Tell Me About Your Latest Prompt"
- 40:30: Sumble: Staying Out of the LLM Kill Zone
- 42:13: Closing Thoughts

## Clips

Short segments published from this episode. Each is citable on its own.

- **LLMs Do the Thinking. The Corpus Does the Grounding.**
  - URL: https://www.youtube.com/watch?v=apwxgKCvH5E
  - Speaker: Anthony Goldbloom
  - Occurs at: 20:21
  - Duration: 0:12
  - Claim: LLMs think, the corpus grounds
- **AI Agents Can 10X Your Outreach. A Human Still Designs It**
  - URL: https://www.youtube.com/watch?v=1pcBUeYq9hM
  - Speaker: Carol Carpenter
  - Occurs at: 8:31
  - Duration: 0:34
  - Claim: Agents scale outreach, humans design it
- **Greg Brockman Learned AI on Kaggle: Why Data Beats Models**
  - URL: https://www.youtube.com/watch?v=_VGUnbTUudM
  - Speaker: Anthony Goldbloom
  - Occurs at: 4:23
  - Duration: 0:23
  - Claim: Data beat models, so he built Sumble
- **Claude and ChatGPT Can't Prioritize Your Funnel**
  - URL: https://www.youtube.com/watch?v=SqDU3bFFrYU
  - Speaker: Carol Carpenter
  - Occurs at: 9:04
  - Duration: 0:39
  - Claim: 57 touch points became 152
- **Claude Is Good at SQL: You Don't Need a Knowledge Graph**
  - URL: https://www.youtube.com/watch?v=lfJKhZGH7TQ
  - Speaker: Anthony Goldbloom
  - Occurs at: 29:07
  - Duration: 0:40
  - Claim: A warehouse gets you 70% of the value
- **How an AI Data Company Uses AI**
  - URL: https://www.youtube.com/watch?v=OK8jpGISwuE
  - Speaker: Carol Carpenter
  - Occurs at: 2:56
  - Duration: 0:49
  - Claim: Three ways Cohesity uses AI
- **Building AI Is 80% Cleaning Data**
  - URL: https://www.youtube.com/watch?v=_MaS1lrbia8
  - Speaker: Anthony Goldbloom
  - Occurs at: 24:13
  - Duration: 0:17
  - Claim: 80% cleaning, 20% complaining
- **AI Agents for Salespeople**
  - URL: https://www.youtube.com/watch?v=s9TqgqtOIA0
  - Speaker: Carol Carpenter
  - Occurs at: 28:06
  - Duration: 0:40
  - Claim: AI agents for salespeople
- **The LLM Kill Zone: Kaggle Founder on Staying Out of It**
  - URL: https://www.youtube.com/watch?v=EXNmrV7Es2w
  - Speaker: Anthony Goldbloom
  - Occurs at: 41:08
  - Duration: 0:28
  - Claim: Staying out of the LLM kill zone

Carol Carpenter is Chief Marketing Officer of Cohesity. Anthony Goldbloom is CEO and co-founder of Sumble, and before that co-founder and CEO of Kaggle. One runs a marketing organization that buys, builds and uses AI daily. The other builds the data layer underneath it. The conversation is about what has paid back so far, and what the agents still cannot do without a person.

## The number that reframes the problem

Carol opens with a figure that explains why marketing is a data problem now:

> "10 years ago, we used to say you have to touch a customer 57 times in enterprise B2B. I just heard the other day, it's 152 touch points."

Those touches are spread across a website visit, a webinar, a Gartner enquiry, G2, Reddit. Finding and sequencing them is what AI is being asked to do.

## Where the returns are real

The translation work is the clearest win. Cohesity does business in 140 countries with 12 prioritized languages, and the cost of translation is significant. The team built on OpenAI, loading in brand guidelines, tone, documentation standards, cultural nuance, and a memory of everything already translated. The goal is a 60% reduction in cost, and Carol believes they have already shown it is reachable.

Further down the funnel, SDR follow-up went from touching about 20% of marketing qualified leads to close to 80%. Carol is specific that this was not solved by hiring 200 more SDRs; it was a few more people plus the tooling.

## Where it has not delivered

> "Do not underestimate, we underestimated the level of work. It is far more painful than I anticipated."

Carol names two areas where AI is not yet a game changer. Image generation still reads as AI-generated. And the automated customer journey still requires a person to think through why someone buys and what the moments of truth are, before an agent has anything to learn from.

## Why the data matters more than the model

Anthony's argument is that the reasoning is the easy part now:

> "You could think of the LLM as quite good for the thinking portion, but then you need to apply the thinking to a hard-grounded corpus that has the world's companies in it, and knowledge of what the world's companies are doing."

Sumble's answer is a knowledge graph built at the team level. A conventional data vendor tells you a bank uses a particular backup product. Sumble tells you which teams inside it use what, and which of those teams are in a buying window because their architecture is changing.

The unglamorous half of that is data cleaning, which Anthony puts at 80% of the work, including keeping up with job titles that drift from SalesOps to RevOps to GTM engineering.

## The disagreement worth listening to

Aparna pushes on whether Cohesity should build its own knowledge graph over its first-party data. Anthony's answer is that a warehouse gets most of the way there, because language models are good at SQL, and that a private graph is still too manual to be worth it. Carol's position on the underlying question is not negotiable:

> "There's all this proprietary data that we would never ship externally to share, even if it would make some agents smarter."