Anthony Goldbloom co-founded Kaggle and later co-founded Sumble with Ben Hamner. Sumble builds company, team, people, and technology data for account research, scoring, enrichment, and sales signals. Its current product direction also packages workflows as agent skills. The useful comparison with Kaggle is not that the products are identical, but that both organize scattered evidence so users can make a decision.

The verified Kaggle connection

Google Cloud’s 2017 event summary welcomed Kaggle as an acquisition, describing it as a community of data scientists and machine-learning practitioners. This establishes the acquisition and category. It does not support later user totals, transaction terms, or claims about Goldbloom’s private reasons for leaving Google.

Sumble’s about page identifies Goldbloom and Hamner and explains its use of public, traceable evidence for commercial intelligence. The company’s own history should be treated as a primary account, not independent verification of market position.

What Sumble sells

The Sumble documentation describes a web application and enterprise services. Its enterprise-services overview covers enrichment and signals delivered through systems such as Snowflake, Salesforce, Databricks, storage, and email.

The core value proposition is “why now” evidence: changes in teams, hiring, technology use, and projects that may affect whether an account is relevant. A signal is not proof of purchase intent. Job posts can reflect exploration, replacement hiring, or a requirement that never reaches production. Good workflows preserve the date and source behind a signal and let a seller distinguish confirmed adoption from inference.

From web application to agent workflow

Goldbloom’s current account-research guide describes a skill that can combine Sumble data with a user’s CRM, call notes, and other connected context to produce an account brief. The guide explicitly separates internal evidence, Sumble fields, and cited web sources.

This agent-oriented design can reduce manual research. It also raises control questions. A buyer should know which systems the agent can read, which steps spend credits, whether it can write to a CRM or sequencer, how it handles stale identities, and what requires confirmation.

Scoring should remain explainable

Sumble’s account-scoring guide argues for visible factors and warns that incomplete attributes can distort rankings. That is a sound design principle. A score should prioritize review, not declare that an account will buy.

Teams should calibrate weights against their own wins and losses, test stability across territories, and inspect false positives. Firmographic fit, relationship history, and recent triggers should remain separate enough to explain why a rank changed.

Evidence limits on scale and economics

The available public sources do not reliably establish an exact current customer count, conversion rate, annual contract value, revenue growth rate, knowledge-graph size, or price schedule. These figures should not be treated as fact without a dated primary filing or company statement that defines the measure. Funding, valuation, and usage figures also answer different questions and should not be combined into a single claim of traction.

The product is a data-and-provenance layer

Sumble’s account, person, team, technology, and project records are valuable only if a user can distinguish observation from inference. A company job post can support “the company advertised for this skill on this date.” It cannot by itself support “the company has deployed this technology” or “the buyer has budget.” The product should retain the raw source, observation date, extraction method, and confidence behind the normalized field.

Record typeUseful decisionError to avoidVerification step
Firmographicinitial territory or segment fitstale size or ownershipcompare a current company or filing record
People and reporting linecandidate contact mapidentity collision or departed employeeopen the dated source and current profile
Technology signalintegration or displacement hypothesismention mistaken for production useseek a first-party technical or procurement trace
Hiring or project signaltiming hypothesisjob post mistaken for approved initiativecorroborate with company statements and conversation
Relationship historynext-action contextduplicate CRM recordsreconcile account, contact, and opportunity IDs

This provenance layer is more important in an agent workflow because generated prose can turn a weak signal into a confident recommendation. Every number and material claim in a brief should be traceable back to a dated field, internal record, or cited web source.

Entity resolution is the hidden technical challenge

Sales systems contain subsidiaries, parent companies, brands, domains, regional entities, duplicate people, former employees, and shared names. Enrichment that attaches a correct fact to the wrong entity is worse than a missing field because it can look authoritative.

Buyers should test a labeled sample with hard cases: acquired companies, rebrands, multinational subsidiaries, stealth companies, common names, job changes, and contacts with several affiliations. Report precision and coverage separately. A system can achieve broad coverage by making more uncertain matches. The right threshold depends on the downstream action.

Corrections should propagate. If a user merges accounts or rejects a person match, determine whether the fix remains in the CRM, Sumble record, derived score, and later agent brief. Measure how long source changes take to appear and whether historical evidence remains available for audit.

Fit, trigger, and relationship should stay separate

An account can fit the ideal customer profile without a current reason to buy. A strong trigger can occur at a poor-fit account. A warm relationship can justify contact without changing either. Keeping these dimensions separate makes the score explainable and lets sales leadership change strategy without retraining an opaque model.

Calibrate weights on an outcome such as a qualified meeting, opportunity progression, or retained customer, not email opens alone. Use a time split so a score is tested on events that occurred after the calibration period. Compare it with a simple baseline and review both false positives and false negatives by segment and territory.

Do not train only on historical wins. Sales coverage and past qualification rules determine which accounts received attention, so missing outcomes are not random. Qualitative review and controlled prospecting can test whether the model is repeating old focus rather than identifying new demand.

Agent permissions and side effects

The account-research guide says internal context should outrank external data, that evidence should be cited, and that the workflow keeps users involved before spending credits or sending. Those are first-party design claims. Buyers should verify them in the installed skill, connected systems, and product version.

Use a permission ladder: read approved sources; reveal paid contact data only after confirmation; draft but do not send; write to a sandbox or review queue; then allow bounded production writes. Each tool call should record the actor, account, inputs, effect, cost, and result. Idempotency matters when a CRM update succeeds but the agent retries after a timeout.

The voluntary NIST AI Risk Management Framework provides a useful govern-map-measure-manage cycle for an agent using company and personal data. It does not certify Sumble or the connected chat platform. The deployment owner must still test prompt injection, excess permissions, sensitive-data exposure, hallucinated contacts, and recovery from partial actions.

Privacy, lawful use, and outreach quality

Public availability does not make every data use appropriate or accurate. Buyers should identify data categories, sources, collection purpose, legal basis where applicable, retention, correction, deletion, and opt-out handling. Contact details and inferred professional interests deserve tighter access than aggregate company attributes.

An agent should not cite surveillance-like details in outbound copy. Sumble’s guide makes a similar distinction by using evidence to form a relevant hypothesis without telling a prospect exactly where it was observed. That can improve tone, but it does not remove privacy or accuracy obligations. The seller remains responsible for the claim and the message.

A buyer evaluation

Build a representative gold set of accounts and contacts. Measure field-level precision, coverage, freshness, entity matching, provenance availability, and correction behavior. Then test a real workflow against the current process: research time, qualified-account yield, seller acceptance, corrected fields, meetings, opportunity progression, credit cost, and complaints.

Run the agent in read-only mode first. Inspect every source and recommendation, introduce stale and conflicting records, and confirm that uncertainty remains visible. Only grant CRM or sequencer writes after permissions, approvals, receipts, and rollback have passed an end-to-end exercise.

Goldbloom’s Kaggle history supports the claim that he has experience building a data and machine-learning community. It does not establish Sumble’s product accuracy or commercial results. Those should be judged through current field-level evidence and controlled buyer measurements.

Bottom line

Goldbloom’s second company applies structured data and explainable ranking to go-to-market work, now with an explicit agent layer. Sumble should be evaluated on data coverage, provenance, freshness, identity accuracy, integration behavior, user control, and measured lift in qualified work. Founder reputation and unsourced growth figures are not substitutes for that evidence.