6 Best Conversational AI Tools for MMM and Incrementality Testing in 2026: An In-Depth Comparison

85 min read
Published May 26, 2026
Updated

Research disclosure: Sellforte designed and published this evaluation and is one of the vendors assessed. Scores reflect how Claude and ChatGPT interpreted public sources under the stated methodology; this research was not a hands-on product test. A score of zero means no public evidence was found under that methodology, not that a capability is absent. Vendors may submit documentation for re-evaluation at research@sellforte.com. Buyers should verify current capabilities directly with vendors.

Summary: Conversational AI tools for MMM and incrementality testing

This research compares Sellforte, Triple Whale, Lifesight, Fospha, Mutinex, and Prescient AI across 48 criteria in nine categories. Claude and ChatGPT each scored the vendors using publicly available information in September 2026.

Sellforte received the highest average total score at 39.8 out of 48, followed by Triple Whale at 34.0 and Lifesight at 32.9. Results differed by category: Triple Whale received the highest assigned score for agentic execution and autonomy, while Sellforte received the highest assigned scores for channel-level optimization, campaign and ad set-level optimization, and incrementality testing.

The table shows the average scores assigned by Claude and ChatGPT, with the maximum available points for each category. Use the criteria relevant to your team to guide questions for vendors, and see the evaluation and scoring methodology for evidence standards and broader-platform credit rules.

Research scores by vendor and category

Average scores assigned by Claude and ChatGPT, rounded to one decimal. Blue shading shows the proportion of available points received; green outlines identify the highest research score in each category, including ties. Shading and highest-score comparisons use unrounded scores.

Research category Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI
Total research score / 48 39.8 34.0 32.9 26.9 24.3 22.6
1. Marketing data reporting / 4 3.5 3.8 3.9 (highest research score in this category) 2.6 3.9 (highest research score in this category) 3.9 (highest research score in this category)
2. Historical performance insights / 5 5.0 (highest research score in this category) 3.9 3.9 4.1 4.3 4.8
3. Channel-level optimization / 5 4.6 (highest research score in this category) 3.6 4.5 2.4 4.5 3.4
4. Campaign & ad set optimization / 6 4.3 (highest research score in this category) 3.5 2.9 2.8 1.1 2.1
5. Incrementality testing / 5 3.8 (highest research score in this category) 2.9 2.6 1.4 0.6 0.6
6. Agentic execution & autonomy / 4 2.1 3.9 (highest research score in this category) 3.1 2.4 0.1 0.4
7. Conversational interface / 8 6.8 (highest research score in this category) 6.8 (highest research score in this category) 4.0 5.8 3.5 3.5
8. Analytical backbone / 4 3.5 3.8 (highest research score in this category) 3.0 2.5 2.8 3.8 (highest research score in this category)
9. Enterprise-grade platform / 7 6.3 (highest research score in this category) 2.0 5.0 3.0 3.5 0.3

Scroll horizontally to compare all six vendors.

Research takeaways by vendor

These takeaways describe the average scores assigned by Claude and ChatGPT within this six-vendor comparison.

  • Sellforte — 39.8 / 48. Received the highest total score and the highest assigned scores in five categories: historical performance insights, channel-level optimization, campaign and ad set-level optimization, incrementality testing, and enterprise-grade platform.

  • Triple Whale — 34.0 / 48. Received the second-highest total score and the highest assigned score for agentic execution and autonomy (3.9 / 4).

  • Lifesight — 32.9 / 48. Tied for the highest assigned score in marketing data reporting (3.9 / 4) and tied for the second-highest score in channel-level optimization (4.5 / 5).

  • Fospha — 26.9 / 48. Its highest-scoring category, as a proportion of available points, was historical performance insights and causal explanation (4.1 / 5).

  • Mutinex — 24.3 / 48. Tied for the highest assigned score in marketing data reporting (3.9 / 4) and tied for the second-highest score in channel-level optimization (4.5 / 5).

  • Prescient AI — 22.6 / 48. Received the second-highest assigned score for historical performance insights and causal explanation (4.8 / 5), and tied for the highest analytical-backbone score (3.8 / 4).

About the research and table of contents

The 48 evaluation criteria were derived from 1,660 real marketer prompts and more than 700 discussions with marketers and marketing analytics professionals. They cover reporting, performance insights, optimization, incrementality testing, agentic execution, the conversational interface, the analytical backbone, and enterprise requirements.

Sellforte published the criteria, scoring instructions, and research records so readers can inspect how Claude and ChatGPT assigned scores and repeat the assessment. The methodology explains the scope, assumptions, and limitations of the research.

Explore the results, methodology, and vendor sections below:

What are Conversational AI tools for MMM and Incrementality Testing?

Conversational AI tools for MMM and Incrementality Testing help marketers measure media performance, optimize marketing spend allocation and execute optimization actions through a conversational, natural language AI interface.

Unlike traditional dashboards, which require users to navigate filters and find the right charts, AI tools for MMM and Incrementality Testing let marketers interact with their measurement stack the way they would with a senior analyst: "Why did revenue drop last week?", "What's the incremental ROAS of Meta vs. TikTok?", "What happens if I cut re-allocate 20% from retargeting to prospecting?".

Unlike generic LLM tools (ChatGPT, Claude, Gemini), AI tools for MMM and Incrementality Testing are purpose-built for marketing, grounded on advertiser's actual data and real incrementality-based measurement models.

To illustrate a conversational AI tool for MMM and Incrementality Testing, below is a screenshot from Sellforte:

image-png-Oct-31-2025-07-46-29-7178-AM

Categories of conversational AI tools for MMM and Incrementality Testing: What types of tools exist?

There's three categories of conversational AI tools for MMM and Incrementality Testing, from basic to advanced

1. Conversational AI tools for Marketing Data Reporting provide marketers fast access to raw data from advertising platforms, ecommerce platforms or analytics tools. The tools in this category are not connected to an MMM or incrementality testing. These types of tools are commonly found from for example from data connector companies. Since they lack the analytical backbone for measuring the true incremental impact of advertising, their relevance for optimizing media spend allocation is limited.

2. Conversational AI tools for channel-level optimization help marketers optimize channel-level media spend allocation. Marketers can ask questions, such as "What's an optimal budget allocation for the next quarter?". These types of tools are commonly found from traditional Marketing Mix Modeling companies, who have started building AI on top of their MMM. While highly useful for channel-level spend optimization, these tools have a limitation: they are not able to offer insights on the campaign & ad set level where the practical execution of media spend optimization happens. These means that they are also not able to offer autonomous agents for executing media spend optimization actions.

3. Agentic MMM and Incrementality testing tools for Campaign & Ad set level Execution help marketers measure marketing ROI, optimize marketing spend allocation, and execute bidding changes on advertising platforms. This is technologically the most advanced AI tool category. These tools operate on the campaign & ad set level, enabling marketers execute bidding changes on ad platforms. These platforms provide AI Agents for autonomous media spend optimization, based on true incrementality. Agentic MMM tools are part of this category.

How do Conversational AI tools for MMM and Incrementality Testing work?

Conversational AI tools for MMM and Incrementality Testing are an additional layer in an existing MMM and Incrementality Testing technology stack, leveraging all the other layers in the technology stack to answer marketer's questions.

Here's a few examples:

  • For basic data questions ("What was the spend on Google Performance Max last month?"), the AI connects with the Processed Data layer
  • For questions on past performance ("What was the ROI for Google Performance Max last month?"), the AI connect with the data science layer, which has MMM's historical performance results
  • For spend optimization questions ("What's the optimal spend allocation across channels for the next month?") , the AI connects with dynamic tools in the Optimization tools -layer

The tech stack of a conversational AI tool for MMM and Incrementality Testing is illustrated in the image below, using Sellforte as an example:

Tech stack of an AI tool for MMM and Incrementality Testing

Research Methodology: Evaluation Criteria

Each Conversational AI tool for MMM and Incrementality Testing in this article is evaluated against criteria that we formed through extensive primary research based on three lenses. These criteria were developed and this evaluation was authored by Sellforte, one of the platforms included in the comparison.

1. Actual AI tool usage today based on real prompts. We first investigated how marketers are actually using AI tools for MMM and Incrementality Testing today. We collected 1,660 prompts recently made in Sellforte's conversational AI tool, Sellforte AI, by marketers and marketing analytics professionals, and categorized them based on the topic and specific use-case. We included all prompts, independent of whether Sellforte AI was capable of answering marketers.

2. Communicated expectations towards conversational AI tools. We investigated the expectations that marketers and marketing analytics teams have for AI tools for MMM and Incrementality Testing. We analyzed AI-related comments and statements in more than 700 discussions with Sellforte customers and prospects, including marketers, marketing analytics leads, and data scientists working in advertising-heavy industries such as retail, ecommerce, DTC, travel & hospitality, and restaurants. We also reviewed AI requirements in recent MMM and Incrementality Testing RFP documentations.

3. Existing capabilities in Conversational AI tools for MMM and Incrementality Testing in the market. We reviewed 78 vendors operating in the marketing measurement space. We evaluated whether the vendors had an AI offering, and if so, we synthesized the capabilities that they communicate on the website, technical documentation or demos. 

The result: 48 evaluation criteria across 9 categories.

Evaluation criteria for Conversational AI tools for MMM and Incrementality Testing

This table summarizes our evaluation criteria for AI tools for MMM and Incrementality Testing:

Evaluation criteria for Conversational AI tools for MMM and Incrementality Testing Criteria were formed based on extensive primary research: Real AI prompts, Customer interviews, AI Product research
ID Category Criterion What it means
1. Marketing data reporting via conversational AI
1.1 Marketing data reporting via conversational AI AI reports sales progress for online sales AI can report and summarize online sales development.
1.2 Marketing data reporting via conversational AI AI reports sales progress for offline store sales AI can report and summarize sales development in offline store sales, using offline store sales data from customer's data warehouse.
1.3 Marketing data reporting via conversational AI AI reports digital media data (spend, impressions, clicks) AI reports digital media data across Meta, Google, TikTok, and other major paid platforms.
1.4 Marketing data reporting via conversational AI AI reports offline media data AI reports data for offline media, covering at least TV, out-of-home, radio, print.
2. Historical performance insights & causal explanation via conversational AI
2.1 Historical performance insights & causal explanation via conversational AI AI measures incremental ROAS and incremental revenue for each digital channel AI reports true incremental impact, not just last-click or platform-reported ROAS, for each digital channel.
2.2 Historical performance insights & causal explanation via conversational AI AI measures incremental ROAS and incremental revenue for each offline channel AI reports incremental ROAS and revenue for offline channels such as TV, OOH, and radio.
2.3 Historical performance insights & causal explanation via conversational AI AI reports promotion-driven revenue, in addition to media-driven AI surfaces how promotions and pricing changes contributed to sales, not just paid media.
2.4 Historical performance insights & causal explanation via conversational AI AI-reported incremental ROAS is updated daily The AI provides daily measurement of incremental ROAS based on MMM, not just weekly, monthly, or quarterly model refreshes.
2.5 Historical performance insights & causal explanation via conversational AI AI explains why performance has changed AI ties sales changes to specific drivers: promotions, seasonality, weather, media saturation, and more.
3. Channel-level optimization with conversational AI
3.1 Channel-level optimization with conversational AI AI recommends optimal budget allocation by channel AI recommends optimal budget allocation across channels (Meta, Google, TikTok, etc.).
3.2 Channel-level optimization with conversational AI AI forecasts total revenue based on optimal budget allocation AI forecasts expected total revenue under the recommended allocation.
3.3 Channel-level optimization with conversational AI AI provides miROAS and response curves for each channel AI provides Marginal Incremental ROAS (miROAS) and saturation / response curves per channel.
3.4 Channel-level optimization with conversational AI AI supports basic custom scenario planning (e.g., "what if I cut Meta by 20%?") Users can simulate custom what-if scenarios in natural language and see forecasted revenue outcomes.
3.5 Channel-level optimization with conversational AI AI supports advanced scenario planning in natural language (constraints, multi-dimensional optimization…) Users can ask complex scenario questions including constraints, multi-dimensional optimization, and marginal budget recommendations ("where should my next €500K go?").
4. Campaign & ad set-level optimization for Digital Channels with conversational AI
4.1 Campaign & ad set-level optimization for Digital Channels with conversational AI AI provides incremental ROAS of each campaign & ad set AI reports incremental revenue and ROAS at the individual campaign and ad set level, not just at the channel level.
4.2 Campaign & ad set-level optimization for Digital Channels with conversational AI AI provides comparison of incremental ROAS to last-click and ad platform attribution ROAS AI shows the delta between what the ad platform claims and what the model estimates is actually incremental, by campaign and ad set.
4.3 Campaign & ad set-level optimization for Digital Channels with conversational AI AI provides miROAS for each campaign & ad set AI provides Marginal Incremental ROAS at the campaign and ad set level to inform bidding decisions.
4.4 Campaign & ad set-level optimization for Digital Channels with conversational AI AI recommends optimal spend for each campaign & ad set AI recommends optimal spend/budget for each campaign and ad set.
4.5 Campaign & ad set-level optimization for Digital Channels with conversational AI AI recommends optimal bid value for each campaign & ad set AI recommends specific bid values (e.g., Target ROAS) per campaign and ad set for execution on ad platforms.
4.6 Campaign & ad set-level optimization for Digital Channels with conversational AI AI provides pre/post analysis for each bidding change After a bidding or budget change is applied, the AI measures the actual impact and provides a pre/post comparison at the campaign & ad set level.
5. Incrementality testing with conversational AI
5.1 Incrementality testing with conversational AI AI can summarize results for Geo Tests AI generates detailed reports for geo incrementality tests, including iROAS and confidence intervals.
5.2 Incrementality testing with conversational AI AI can summarize results for Own Media A/B tests (e.g. leaflet tests) AI generates detailed reports for A/B tests for own media, such as leaflet tests.
5.3 Incrementality testing with conversational AI The platform ingests Meta Conversion Lift tests, and the AI can summarize their findings The platform ingests Meta Conversion Lift tests, and the AI can generate detailed reports based on the test.
5.4 Incrementality testing with conversational AI AI provides incrementality testing recommendations AI recommends what to prioritize for incrementality testing next (e.g. geo testing, A/B testing, or conversion lift testing).
5.5 Incrementality testing with conversational AI AI provides incrementality test design recommendations AI recommends incrementality test designs, including control group selection, test duration, and statistical power requirements.
6. Agentic execution & autonomy
6.1 Agentic execution & autonomy AI can push daily spend/budget changes to Meta, Google, TikTok APIs The AI can execute daily spend/budget changes directly on major ad platforms via API.
6.2 Agentic execution & autonomy AI can push bidding changes (e.g., Target ROAS) to Meta, Google, TikTok APIs The AI can execute bidding parameter changes (e.g. Target ROAS) directly on major ad platforms via API.
6.3 Agentic execution & autonomy AI's level of autonomy for execution can be configured The platform supports a configurable spectrum: insight-only → recommendation-only → human-approved execution → fully autonomous execution. Users choose per use case.
6.4 Agentic execution & autonomy Proactive insights & alerts via AI AI surfaces anomalies, narratives, and scheduled reports proactively via Slack or email, without the user having to ask.
7. AI's UX & conversational interface
7.1 UX & conversational interface Includes tables and charts inline in AI responses AI inlines data visualizations directly inside chat responses, not just text.
7.2 UX & conversational interface AI has conversation history & multi-turn context retention Follow-up questions retain prior context: dimensions, filters, time windows, and named entities. "What about for the next 8 weeks?" knows what "next" refers to.
7.3 UX & conversational interface Allows convenient export of AI outputs (e.g., PDF, Slides, CSV) AI outputs can be exported to PDF, Slides, or CSV for sharing outside the platform.
7.4 UX & conversational interface AI grounds answers in data, providing links from outputs to deep-dive dashboards AI outputs link out to deep-dive dashboards or detail views for further investigation, connecting the chat to the underlying data.
7.5 UX & conversational interface AI operates in embedded and multi-window mode (chat + dashboard side-by-side) The AI chat can be shown side-by-side with a dashboard, allowing users to submit prompts about what they are viewing.
7.6 UX & conversational interface AI shows its reasoning steps and logic while answering The AI surfaces its reasoning steps, the data sources it pulled from, and the assumptions behind each answer as the answer is being constructed.
7.7 UX & conversational interface AI handles non-English questions in production AI responds accurately in the user's language for at least 5 major languages, maintaining domain accuracy, not just translating.
7.8 UX & conversational interface AI can be accessed via MCP server / external LLM The tool exposes data via MCP server or equivalent, allowing external LLMs (Claude, ChatGPT, Cursor) to query it directly.
8. Analytical backbone of the Conversational AI
8.1 Analytical backbone of the Conversational AI AI provides deterministic, model-backed answers with Bayesian MMM as backbone The AI's recommendations are grounded in a Bayesian Marketing Mix Model, not last-click attribution or descriptive analytics.
8.2 Analytical backbone of the Conversational AI Bayesian MMM used by the AI is calibrated with incrementality tests The MMM is calibrated against real incrementality test results, with priors and posteriors informed by causal experiments.
8.3 Analytical backbone of the Conversational AI AI reports model validation and other modelling KPIs The tool reports model validation metrics (R², MAPE, posterior predictive checks, holdout performance) so users can assess model quality.
8.4 Analytical backbone of the Conversational AI Model calibration & configuration settings (e.g., priors) are auditable and editable in a self-serve UI Customers can inspect and configure model priors and other key parameters in a self-serve UI, not just accept the model as a black box.
9. Enterprise-grade platform
9.1 Enterprise-grade platform At least 10 public reference customers from $1B+ revenue brands Proven track record with large, sophisticated advertisers, not just mid-market or DTC brands.
9.2 Enterprise-grade platform SOC 2, ISO 27001, or audited IT security by a third-party cyber security auditor Independently verified security posture, a baseline requirement for enterprise IT procurement.
9.3 Enterprise-grade platform Data residency: geography option between US and EU Customers can choose where their data is stored, critical for GDPR compliance in Europe.
9.4 Enterprise-grade platform Multi-cloud: option between AWS, GCP, and Azure Deployment flexibility to match the customer's existing cloud infrastructure.
9.5 Enterprise-grade platform Supports single sign-on (SSO) for enterprises Enterprise authentication via SSO, required by most large-company IT policies.
9.6 Enterprise-grade platform Customer data is not shared to a third-party LLM (LLM is deployed in a customer-specific cloud container) AI inference runs within isolated cloud infrastructure; customer data does not egress to third-party LLM APIs such as OpenAI.
9.7 Enterprise-grade platform Hands-on demo or trial of the AI is available without sales-call gating A real, clickable demo of the AI is publicly accessible without requiring a sales call, a signal of product confidence and evaluation friendliness.

Let's next cover briefly what each of these categories mean and why they matter.

Category 1. Marketing data reporting via conversational AI

This category measures AI tools' capabilities to provide basic marketing data reporting.

Summarizing historical data is the foundation for any marketing measurement AI tool for MMM and Incrementality Testing. At the simplest level, this means surfacing raw data from ad platforms and ecommerce platforms, such as clicks, impressions, conversions, and ecommerce sales. Enterprise-grade platforms also cover offline store sales and offline media, which might require a data warehouse connection in the backend.

To illustrate marketing data reporting in action, below is a chart from Sellforte showing how impressions have developed across ad platforms for the last 8 weeks.

Sellforte AI reporting basic marketing data: Development of impressions in the last 8 weeks

Category 2. Historical performance insights & causal explanation via conversational AI

This category measures AI tools' capabilities to provide historical performance insights grounded in true incremental sales impact of media, and ability to explain the causal drivers for performance changes.

While some AI tools for MMM and Incrementality Testing are satisfied reporting ROAS from last-click or ad platform attribution, modern AI tools are measuring the true incremental ROAS of each channel and campaign. Most advanced marketing AI tools can also explain the causal drivers behind historical KPI changes, such as changes in base sales, promotion-driven sales, seasonality, or weather.

To illustrate causal explanation in action, below is a screenshot from Sellforte decomposing year-over-year sales decline into its drivers.

Sellforte AI decomposing sales to its drivers

Category 3. Channel-level optimization with conversational AI

This category measures AI tools' capabilities to help marketers optimize media spend allocation across channels. While the two previous categories covered reporting, telling you what happened, we now move to optimization, which tells you what to do next. This is where an AI tool starts becoming truly valuable for marketing teams.

Modern AI tools for MMM and Incrementality Testing can recommend spend reallocation across channels and forecast the revenue impact of recommendations. They can answer "what if I cut Meta by 20% and shift it to YouTube?" in seconds rather than weeks. To do this credibly, the AI needs access to an optimization tool that leverages Marginal Incremental ROAS (miROAS) and Advertising Response Curves for each channel, and can project total revenue under different allocation scenarios. 

The most sophisticated AI tools for MMM and Incrementality Testing go further: handling natural-language constraints, multi-dimensional optimization across channel and geography, and marginal recommendations for "if I got an extra €500K, where should it go?" 

To illustrate channel-level optimization in action, below is a screenshot from Sellforte recommending optimal spend by channel for the next 3 months.

Sellforte AI recommending optimal spend by channel

Category 4. Campaign & Ad set -level optimization for Digital Channels with conversational AI

This category measures AI tools' capabilities to support spend optimization on tactical level: optimizing spend across campaigns and ad sets within digital channels.

Campaign and ad set level is the most important granularity of optimization, because that's where budget execution practically happens. A budget shift from Meta to TikTok at the channel level is meaningless until it's executed as specific budget and bid changes across dozens or hundreds of individual campaigns and ad sets.

Most AI tools for MMM and Incrementality Testing that handle channel-level optimization don't follow through to the campaign layer. The technical bar is higher here, as this level of optimization requires reliable reliable estimation of campaign and ad set level miROAS.

To illustrate campaign and ad set level optimization in action, below is a screenshot from Sellforte recommending Target ROAS bidding changes for specific Google Ads campaigns

Sellforte AI recommending Target ROAS changes

Category 5. Incrementality testing with conversational AI

This category measures AI tools' capabilities to analyze and plan incrementality tests.

Incrementality testing is used in modern measurement to calibrate Marketing Mix Models, as they can provide ground truth for a channel's incremental ROAS at a specific point in time and at a specific spend level. Advanced AI tools for MMM and Incrementality Testing have access to an incrementality test library that contain all incrementality tests done by a company, across geo tests, own media A/B tests and conversion lift tests. They can summarize the main insights from tests, including iROAS and confidence intervals. The most advanced tools also help recommending channels to tests, as well as give guidance for test design.

To illustrate how modern AI tools for MMM and Incrementality Testing can integrate with incrementality tests, below is a screenshot from Sellforte embedding AI into an incrementality testing dashboard providing a summary of how to interpret a test.

Sellforte AI embedded into an incrementality testing dashboard

Category 6. Agentic Execution & Autonomy

This category measures AI tools' capabilities for agentic execution.

The next generation of AI tools for MMM and Incrementality Testing have agents that take action by adjusting bidding parameters directly in ad platforms. This is a fundamentally different product category, and it requires different design choices: clear autonomy boundaries, approval workflows for high-impact actions, audit logs, and rollback capabilities.

To illustrate how modern AI tools for MMM and Incrementality Testing can adjust bidding parameters, below is a screenshot from Sellforte's approval flow for adjusting Target ROAS for a Google Ads campaign.

Sellforte AI approval flow for bidding value change

Category 7. AI's UX & Conversational Interface

This category measures AI tools' user-experience and features available in the conversational chat interface.

Tested features in this category include chat history, context retention, ability to provide inline tables and charts, multi-window mode, and exportable outputs (PDF, Slides, CSV). We also test features building trust, such as ability to provide reasoning flow and links from AI answers to source data and deep-dive dashboards.

To illustrate reasoning flows in action, below is a screenshot from a Sellforte describing its interpretation of a user prompt as well as steps its taking to reply to it.

Sellforte AI's reasoning flow

To illustrate a dual-window mode in action, below is a screenshot from Sellforte where AI interface is next to a budget optimizer tool. The AI on the right side of the screen can be asked to use the optimizer to find optimal budget allocations, and the user can continue scenario configuration in a manual mode on the left side of the screen if needed.

Sellforte AI in dual window mode

Category 8. Analytical Backbone of the Conversational AI

Everything in the previous categories measured what the AI does. This category is about what the AI is built on.

Modern marketing measurement and optimisation systems are built on three core methodologies: Marketing Mix Modeling, incrementality testing and attribution, as illustrated in the chart below. 

Evaluate criteria include whether the AI tool for MMM and Incrementality Testing is powered by a Bayesian Marketing Mix Model calibrated with incrementality tests, whether model validation features are available, and whether there are transparent model configuration and calibration tools available for the user.

Category 9. Enterprise-grade platform

This final category measures whether the AI tools can operate in an enterprise environment, serving large organizations with mature IT policies.

The dimensions scored include large client references, security certifications (SOC 2, ISO 27001), data residency options between US and EU regions, multi-cloud support across AWS, GCP, and Azure, and enterprise authentication via single sign-on.

While mid-sized brands might not yet express all of these requirements, they will ultimately grow into a size where such requirements become not just relevant but mandatory for procurement approval.

What biases & limitations does our evaluation criteria have, and how are we addressing them?

Bias 1: Industry sample bias toward Ecommerce and Retail. The 1,660 prompts and 700+ customer discussions that informed the criteria come predominantly from Sellforte's customer and prospect base, which skews heavily toward Retail and Ecommerce. This means certain use cases that might matter more in other verticals, such as brand measurement in CPG/FMCG may be underrepresented. 

Bias 2: Coverage of requirements from both small and large businesses. The criteria span a wide range, from table-stakes capabilities that even small ecommerce brands need (basic data reporting, channel-level optimization) to enterprise-grade requirements (multi-cloud, SSO, EU data residency, $1B+ customer references). This breadth is intentional but means no single vendor scores perfectly: a tool purpose-built for SMBs will be penalized on enterprise criteria it has no reason to meet, and vice versa. We address this partially through the "Best for" summaries for each vendor, which contextualize the scores against the buyer segment each tool is designed for.

Research Methodology: Scoring

Scoring each tool against the evaluation criteria was done by Claude and ChatGPT. See specific LLM-model versions in Evaluation Dates by Vendor and LLM. Here were the assessment steps:

Step 1. Both LLMs independently scored each criterion for each vendor, based on the instructions in this section. 

Step 2. For the final score for each criteria, we used an average between the two LLMs, rounded to 1 decimal.

Step 3. Category scores were created by summing up the scores from each criteria in the category, and total score was calculated by by summing up the category scores.

Why is the scoring done by Claude and ChatGPT?

Simulating evaluation a buyer might make based on public materials. By using only vendor-provided materials on their website and technical documentation in the assessment, Claude and ChatGPT -based evaluation simulates how a buyer without prior knowledge of the vendor might assess the vendor prior to a sales call or demo meeting.

Scoring by Claude and ChatGPT. Sellforte designed the research and is one of the vendors assessed. Claude and ChatGPT assigned the scores using the published criteria and instructions. Using LLMs does not eliminate potential bias in the research design, source selection, or interpretation of evidence.

Transparent methodology and reproducible assessment. Because the evaluation instructions are published alongside this article, any reader can re-run the evaluation for any vendor and verify or challenge the results. This is a higher standard of transparency than conventional analyst-style research, where the scoring rationale is typically not disclosed.

Public documentation as the evidence base. Claude and ChatGPT were instructed to score vendors using publicly available documentation. The availability and discoverability of that documentation can affect the assigned scores.

Equal treatment of vendors. Claude and ChatGPT apply the same instructions, the same criteria, the same scoring scale, and the same source prioritization rules to every vendor in the comparison.

Public-source research by Claude and ChatGPT. Claude and ChatGPT used web search to review vendor websites and technical documentation. Either model may miss relevant material or interpret it incorrectly, so the research records include the sources and rationales for readers to inspect.

Instructions given to Claude and ChatGPT

The following table reproduces the instructions given to Claude and ChatGPT for each vendor evaluation. Role descriptions in this historical prompt refer to the AI models. These instructions are also available as a Google Sheet: Evaluation instructions: Conversational AI Tools for MMM and Incrementality Testing

Instructions given to Claude and ChatGPT for vendor evaluation
Instruction ID Instruction
1 Act as an independent evaluator of Conversational AI tools for Marketing Mix Modeling and Incrementality Testing.
2 Use evaluation criteria from the Criteria sheet.
3 In the evaluation, only use information available at the company website, and in the domain where technical documentation is located (if separately hosted).
4 When scoring categories 1–6, use the scoring model from the Scoring — Cat 1–6 sheet.
5 When scoring categories 7–9, use the scoring model from the Scoring — Cat 7–9 sheet.
6 When scoring category 1, you can assume a score of 1 if you find evidence of the overall platform supporting the capability.
7 When scoring criteria 2.1, 2.2, and 2.3, you can assume a score of 1 if you find evidence of the overall platform supporting the capability.
8 If you find evidence that the conversational AI supports optimization of offline media, you can assume that the tool also measures offline media iROAS.
9 If you find evidence that the conversational AI supports optimization of digital media, you can assume that the tool also measures digital media iROAS.
10 When interpreting terminology, you can assume that
- ROI is the same thing as incremental ROAS
- Marginal ROI is the same thing as Marginal Incremental ROAS or miROAS
- Diminishing return curves are the same thing as response curves
11 A vendor's conversational AI (a first-party chat the vendor built into its product) is a separate surface from an MCP server. MCP availability proves the underlying capability exists, so it can support the "broader platform" scoring, but it never earns the top "in conversational AI chat" tiers (1 / 0.75) unless the vendor demonstrates the capability inside its own first-party AI chat.
12 Prioritize source materials in this order: 1. Technical documentation (such as a support center); 2. Product page; 3. Marketing collateral (such as product launch blog posts).
13 As an output, provide your evaluation in an excel file, using these columns:
- Category Number 
- Category ID 
- Criteria Score 
- Two-sentence rationale for the score. If you found no evidence, comment that you did not find evidence, instead of claiming that the capability does not exist
- URL to source 
- Type of source

Let's next discuss the instructions in detail.

Scoring model for Categories 1–6 (Instruction 4)

In instruction 4, we asked the LLMs to use the scoring table below for categories 1 through 6. 

Scoring model for Categories 1–6: Functional AI capabilities
Score Definition
1.0 Strong public evidence that the capability exists in the conversational AI chat interface
0.75 Partial public evidence that the capability exists in the conversational AI chat interface
0.5 Feature is not available via the conversational chat interface, but there is strong public evidence that the capability is available in the broader platform
0.25 Feature is not available via the conversational chat interface, but there is partial public evidence the capability is available in the broader platform
0.0 No evidence found that the feature exists in the conversational AI tool or in the broader platform

Categories 1 through 6 cover functional capabilities of the conversational AI. For these categories, the scoring model distinguishes between capabilities demonstrated inside the vendor's conversational AI chat interface versus capabilities only evidenced in the broader platform. 

We are instructing the LLMs to assign scores 1 or 0.75 if there is evidence (strong or partial) that a feature is available in the Conversational AI interface. 

We asked the LLMs to give a score of 0.5 or 0.25 if there is a capability that exists in the platform, but there is no evidence for it being available via the Conversational AI interface. As an example, a platform might be able to show geo test results, but there is no documentation where the AI can summarize their findings. Why did we apply this instruction? We wanted each of the assessed platform get some credit in these cases as well for two reasons. Firstly, it is likely that the feature will be added to conversational AI in the near future. Secondly, it is possible that the feature can already be accessed via the conversational AI, but the documentation does not yet exist.

A score of 0 means the LLMs could not verify a capability from public sources, not that the capability does not exist.

Scoring model for Categories 7–9 (Instruction 5)

In instruction 5, we asked the LLMs to use the scoring table below for categories 7 through 9. 

Scoring model for Categories 7–9: UX, Analytical Backbone, and Enterprise Platform
Score Definition
1.0 Strong evidence that the platform supports the capability
0.5 Partial evidence that the platform supports the capability
0.0 No evidence found that the platform supports the capability

Categories 7 through 9 cover the conversational interface UX, the analytical backbone of the AI, and the enterprise platform maturity.

For these categories, we instructed LLMs to assign 1 if there is strong evidence that the platform supports the capability and 0.5 if there the evidence is partial but not strong.

A score of 0 means the LLMs could not verify a capability from public sources, not that the capability does not exist.

Scoring Instructions 3 and 12: Source material

Instruction 3: "In the evaluation, only use information available at the company website, and in the domain where technical documentation is located (if separately hosted)."

This constraint keeps the playing field level. Every vendor controls the source material used for assessing their capabilities, and thus every vendor has the same opportunity to influence their scores.

Instruction 12: "Prioritize type of source materials in this order: 1. Technical documentation (such as a support center); 2. Product page; 3. Marketing collateral (such as product launch blog posts)."

Technical documentation is prioritized because it is typically the most precise and the least subject to marketing inflation. Product pages on vendors' website are second, as they also tend to be feature-driven. Marketing collateral, such product launch blog posts are treated as third type of evidence.

Scoring Instructions 6, 7, 8, 9: Assuming basic features

We noticed that vendors' documentation is sometimes focused on the most advanced features of the tools, and the basic features might not be fully spelled out. This is the case especially with vendors who have relatively little focus on public documentation. For this reason, we decided to ask the LLMs to make assumptions that some of the basic features, such as data reporting, exist in the conversational AI, if there's evidence for the overall platform supporting them.

Instruction 6: "When scoring category 1, you can assume a score of 1 if you find evidence of the overall platform supporting the capability."

Instruction 7: "When scoring criteria 2.1, 2.2, and 2.3, you can assume a score of 1 if you find evidence of the overall platform supporting the capability."

The criteria referenced in instructions 6 and 7 cover the most basic data reporting capabilities, such as does the Conversational AI give access to digital media data, or does it report the iROAS / ROI of digital channels. We asked the LLMs to assume they exist in the conversational AI if there was evidence that that broader platform covers it.

Instruction 8: "If you find evidence that the conversational AI supports optimization of offline media, you can assume that the tool also measures offline media iROAS"

Instruction 9: "If you find evidence that the conversational AI supports optimization of digital media, you can assume that the tool also measures digital media iROAS"

These instructions are also related to the existence of basic features. If the LLMs found evidence that digital media spend could be optimized with AI, we asked them to assume that basic measurement of digital iROAS / ROI also existed.

Disclaimer: Instructions 6, 7, 8 and 9 benefit most vendors whose documentation is limited, and is less beneficial for vendors with extensive documentation.

Scoring Instruction 10: Terminology

Instruction 10: "When interpreting terminology, you can assume that
- ROI is the same thing as incremental ROAS
- Marginal ROI is the same thing as Marginal Incremental ROAS or miROAS
- Diminishing return curves are the same thing as response curves"

Different vendors use different terminology for the same underlying concepts. Penalizing a vendor for calling it "Marginal ROI" instead of "miROAS" would introduce scoring noise unrelated to actual capability differences. These equivalences make terminology interpretation consistent across all vendors.

Scoring Instruction 11: Focus on Conversational AI

Instruction 11: "A vendor's conversational AI (a first-party chat the vendor built into its product) is a separate surface from an MCP server. MCP availability proves the underlying capability exists, so it can support the "broader platform" scoring, but it never earns the top "in conversational AI chat" tiers (1 / 0.75) unless the vendor demonstrates the capability inside its own first-party AI chat."

This is one of the most important methodological choices in the evaluation. An MCP server lets an external LLM query a vendor's data, but that is not the same as the vendor's own AI chat interface answering a question. A marketer asking "what was my incremental ROAS on Meta last month?" in a vendor's first-party AI is having a different product experience than a developer who has wired up an external Claude instance to the same vendor's MCP. Both matter, but they are distinct product surfaces, and this evaluation focuses on the first-party conversational AI experience. We might do a separate evaluation of MCPs, as the space develops.

How to interpret the scores?

The scoring framework that the LLMs were asked to use is specifically targeted to evaluate availability of public evidence for the capabilities, and thus it should be interpreted as such. There might be a gap between actual features of the platform and what's publicly documented by the vendors.

How to reproduce the evaluation and scoring yourself with Claude or ChatGPT?

To reproduce the evaluation for any vendor in this study, you can follow the instructions in this section. They are accurate as of 3rd September 2026. Reproducibility instructions may require updating as model versions change.

Step 1. Download the Evaluation Instructions Google Sheet as an Excel file (or other file format you can attach to an LLM prompt): Evaluation instructions: Conversational AI Tools for MMM and Incrementality Testing. We could not get LLMs to access the Google Sheet directly without a Google Drive integration.

Screenshot of Evaluation Instructions in Google Sheet

Step 2. Open Claude in Incognito mode to disconnect its memory about your previous conversations that might influence the evaluation. To achieve the same in ChatGPT, you need to disable ChatGPT memory.

Screenshot of Claude Incognito mode

Step 3. Choose a model. See specific LLM-model versions we used in sectopm Evaluation Dates by Vendor and LLM
Screenshot of Claude model selection

Step 4. Initiate the prompt: "Evaluate [add vendor name, e.g., Sellforte], based on the instructions in the attached Excel."

Future updates of scoring

This comparison is updated on a rolling basis, with a full refresh targeted at least once per quarter.

What biases & limitations does our scoring approach have, and how are we addressing them?

While LLM-based evaluation reduces vendor bias and promotes equal treatment of vendors and reproducibility, there are four main limitations to be aware of.

1. Availability and of documentation that each vendor has made public. A vendor with extensive and detailed public documentation will naturally score higher than one that keeps product details behind a sales call gate, even if the underlying capabilities are comparable. Each vendor can affect their own scoring through public documentation.

2. LLMs' ability to find public documentation. Even with web search enabled, Claude and ChatGPT may not find every relevant page on a vendor's website. Documentation that is not well-indexed or that sits in obscure subdomains may be missed. We tried to address this by using advanced models from two separate LLMs, and added a specific point to the instructions to search for vendor-provided technical documentation that might sometimes be under a different sub-domain.

3. LLMs' ability to interpret public documentation against the scoring criteria. Matching a product description to a specific criterion requires judgment. We reduce interpretation variance by using advanced models from two separate LLMs, providing detailed criterion descriptions, and providing explicit terminology equivalences (instruction 10.0). However, edge cases may still exist.

4. Limitations of LLM technology, including hallucinations. LLMs are known to occasionally assert things that are not based on facts. To reduce this risk, we used advanced models from two separate LLMs, asked them to provide source URL for their assessment, and asked them to provide a rationale for the score in each criterion.

What tools we evaluated: Conversational AI tools for MMM and Incrementality Testing

We focused this article on conversational AI tools for MMM and Incrementality Testing that meet all three of the following criteria:

  1. Real product: Proven AI tool with a conversational interface for MMM and Incrementality Testing that is offered as a dedicated tool or as part of a broader measurement platform.
  2. Used by recognized advertisers: The tool is used in production by enterprise brands, with at least some publicly verifiable customer references.
  3. Recognized by the industry: The tool has visible market presence, including coverage in industry press, analyst reports, social media discussion, or conference talks.

When searching for tools that would meet these criteria we looked into multiple product categories: Data connector companies, traditional Marketing Mix Modeling providers, next gen MMM vendors, Incrementality testing tools, attribution tools and generic AI tools. We ended up evaluating six AI tools for MMM and Incrementality Testing matching the criteria above: Sellforte, Triple Whale, Mutinex, Lifesight, Fospha and Prescient AI.

Surprisingly, we found out that AI adoption is still low in many companies. As an example, at start of 2026, only 13% of Marketing Mix Modeling vendors have implemented AI.

Below is a selection of solutions often associated with the MMM and Incrementality testing space, and a rationale why they were not evaluated in this research.

MMM tools checked but not scored, including rationale (as of Sep 2026). Research conducted with ChatGPT Sol 5.6
Tool checked Reason not included
Google Meridian Open-source MMM framework. No built-in conversational chat interface was found.
Meta Robyn Open-source MMM library. No built-in conversational chat interface was found.
PyMC-Marketing Open-source MMM library. No first-party conversational chat interface was found; third-party wrappers were not counted.
Analytic Partners No public evidence of an official MCP or an MMM-specific conversational chat interface was found.
Circana Circana describes Liquid Mix as an AI-powered, self-service MMM platform, but we found no public evidence of an official MCP or conversational chat interface.
Ekimetrics No public evidence of an official MCP or an MMM-specific conversational chat interface was found.
Ipsos MMA Ipsos offers conversational AI products in other research areas, but we found no public evidence that they connect to Ipsos MMA's measurement platform.
Recast Recast has an official MCP for external AI assistants. We found no evidence of a first-party conversational chat interface.
Cassandra AI-generated, plain-language insights are documented, but we found no public evidence of an interactive conversational chat interface or official MCP.
Keen Decision Systems AI-driven modeling and planning are documented, but we found no public evidence of a conversational chat interface or official MCP.
Measured Measured has an official MCP that brings its data into external AI assistants, but we found no evidence of its own conversational chat interface.
LiftLab No public evidence of an official MCP or an MMM-specific conversational chat interface was found.
DoubleVerify / Rockerbox DoubleVerify has announced MCP-based access through external AI assistants. We found no public evidence of a first-party chat interface grounded specifically in Rockerbox MMM outputs.
Northbeam Northbeam documents machine-learning-based attribution and MMM, but we found no public evidence of a conversational chat interface or official MCP.

Because this category is changing quickly, we will consider additional vendors in future scoring rounds as qualifying interfaces become publicly documented and broadly available.

In-depth Comparison of Conversational AI tools for MMM and Incrementality Testing

The charts and category-specific tables below break the evaluation into the total research score and nine evaluated categories. The full Claude and ChatGPT assessment, including source material and detailed scoring, is available in this Google Sheet: Evaluation: Conversational AI Tools for MMM and Incrementality Testing.

How to interpret this comparison: Sellforte designed and published this evaluation and is one of the vendors evaluated. Scores are averages assigned by Claude and ChatGPT against predefined criteria using publicly available documentation reviewed as of September 2026; this was not a hands-on product test. A score of 0 means that Claude and ChatGPT found no public evidence for the criterion, not that the capability does not exist. Documentation and products may change, so buyers should verify current capabilities directly with each vendor.

Total research score across all 48 criteria

Total research score by vendorScore across all nine categories (maximum 48)
Sellforte39.8 out of 4839.8 / 48
Triple Whale34.0 out of 4834.0 / 48
Lifesight32.9 out of 4832.9 / 48
Fospha26.9 out of 4826.9 / 48
Mutinex24.3 out of 4824.3 / 48
Prescient AI22.6 out of 4822.6 / 48

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. The total research score combines 48 criteria across nine categories. For functional Categories 1–6, the criteria primarily assess whether capabilities are available through each vendor's first-party conversational AI, subject to the broader-platform partial-credit and stated assumption rules described in the methodology. Category 7 evaluates the conversational interface and related access surfaces, while Categories 8–9 evaluate the analytical backbone supporting the AI and the broader platform's enterprise readiness. Each criterion was scored from 0 to 1 by Claude and ChatGPT, and the scores assigned by Claude and ChatGPT were averaged before the category and total scores were calculated.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (39.8 / 48), Triple Whale (34.0 / 48), Lifesight (32.9 / 48), Fospha (26.9 / 48), Mutinex (24.3 / 48), and Prescient AI (22.6 / 48). The category-level results were less uniform than the total ranking. Sellforte received the highest assigned score in Categories 2, 3, 4, 5, and 9; Triple Whale did so in Category 6; Lifesight, Mutinex, and Prescient AI tied in Category 1; Sellforte and Triple Whale tied in Category 7; and Triple Whale and Prescient AI tied in Category 8. The total therefore combines distinct patterns across the evaluated use cases rather than showing one vendor receiving the highest score in every category.

Research scores by category and vendorCategory scores and total score, based on the average of the Claude and ChatGPT evaluations and displayed to one decimal.
Evaluated category Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Max
1. Marketing Data Reporting via conversational AI 3.5 3.8 3.9 2.6 3.9 3.9 4
2. Historical Performance Insights & Causal Explanation via conversational AI 5.0 3.9 3.9 4.1 4.3 4.8 5
3. Channel-Level Optimization with conversational AI 4.6 3.6 4.5 2.4 4.5 3.4 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI 4.3 3.5 2.9 2.8 1.1 2.1 6
5. Incrementality Testing with conversational AI 3.8 2.9 2.6 1.4 0.6 0.6 5
6. Agentic Execution & Autonomy 2.1 3.9 3.1 2.4 0.1 0.4 4
7. AI's UX & Conversational Interface 6.8 6.8 4.0 5.8 3.5 3.5 8
8. Analytical Backbone of the Conversational AI 3.5 3.8 3.0 2.5 2.8 3.8 4
9. Enterprise-Grade Platform 6.3 2.0 5.0 3.0 3.5 0.3 7
Total score out of 48 39.8 34.0 32.9 26.9 24.3 22.6 48

1. Marketing Data Reporting via conversational AI

Category 1 research score by vendorMarketing Data Reporting via conversational AI (maximum 4)
Lifesight3.9 out of 43.9 / 4
Mutinex3.9 out of 43.9 / 4
Prescient AI3.9 out of 43.9 / 4
Triple Whale3.8 out of 43.8 / 4
Sellforte3.5 out of 43.5 / 4
Fospha2.6 out of 42.6 / 4

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category examines 4 criteria that assess whether the vendor's conversational AI can report and summarize online and offline sales data, digital-media data, and offline-media data.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Lifesight (3.9 / 4), Mutinex (3.9 / 4), Prescient AI (3.9 / 4), Triple Whale (3.8 / 4), Sellforte (3.5 / 4), and Fospha (2.6 / 4). All six vendors received 1.0 for online-sales reporting and digital-media reporting, while the offline-sales and offline-media criteria produced most of the variation in the category scores. In this evaluation, the public documentation was therefore more consistent for digital reporting than for offline reporting.

Criteria-level scores for Category 1Each criterion is scored from 0 to 1; the category maximum is 4. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
1.1 AI reports sales progress for online sales 1.0 1.0 1.0 1.0 1.0 1.0 1.0
1.2 AI reports sales progress for offline store sales 0.8 0.9 1.0 0.4 0.9 1.0 0.8
1.3 AI reports digital media data (spend, impressions, clicks) 1.0 1.0 1.0 1.0 1.0 1.0 1.0
1.4 AI reports offline media data 0.8 0.9 0.9 0.3 1.0 0.9 0.8
Category total out of 4 3.5 3.8 3.9 2.6 3.9 3.9 3.6

2. Historical Performance Insights & Causal Explanation via conversational AI

Category 2 research score by vendorHistorical Performance Insights & Causal Explanation via conversational AI (maximum 5)
Sellforte5.0 out of 55.0 / 5
Prescient AI4.8 out of 54.8 / 5
Mutinex4.3 out of 54.3 / 5
Fospha4.1 out of 54.1 / 5
Triple Whale3.9 out of 53.9 / 5
Lifesight3.9 out of 53.9 / 5

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category examines 5 criteria that assess whether the vendor's conversational AI can surface incremental revenue and ROAS for digital and offline channels, report promotion-driven revenue and daily MMM-based measurement updates, and explain changes in performance.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (5.0 / 5), Prescient AI (4.8 / 5), Mutinex (4.3 / 5), Fospha (4.1 / 5), Triple Whale (3.9 / 5), and Lifesight (3.9 / 5). All vendors received 1.0 for reporting incremental ROAS and revenue for digital channels, and the scores were also closely grouped for offline-channel measurement and explanations of performance changes. The largest separation came from the criterion covering daily MMM-based incremental ROAS updates, followed by documentation of promotion-driven revenue.

Criteria-level scores for Category 2Each criterion is scored from 0 to 1; the category maximum is 5. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
2.1 AI measures incremental ROAS and incremental revenue for each digital channel 1.0 1.0 1.0 1.0 1.0 1.0 1.0
2.2 AI measures incremental ROAS and incremental revenue for each offline channel 1.0 1.0 0.9 0.9 1.0 1.0 1.0
2.3 AI reports promotion-driven revenue, in addition to media-driven 1.0 0.9 1.0 0.3 1.0 1.0 0.9
2.4 AI-reported incremental ROAS is updated daily 1.0 0.1 0.1 1.0 0.3 0.8 0.5
2.5 AI explains why performance has changed 1.0 0.9 0.9 1.0 1.0 1.0 1.0
Category total out of 5 5.0 3.9 3.9 4.1 4.3 4.8 4.3

3. Channel-Level Optimization with conversational AI

Category 3 research score by vendorChannel-Level Optimization with conversational AI (maximum 5)
Sellforte4.6 out of 54.6 / 5
Lifesight4.5 out of 54.5 / 5
Mutinex4.5 out of 54.5 / 5
Triple Whale3.6 out of 53.6 / 5
Prescient AI3.4 out of 53.4 / 5
Fospha2.4 out of 52.4 / 5

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category examines 5 criteria that assess whether the vendor's conversational AI can provide channel budget allocations, revenue forecasts, marginal incremental ROAS and response curves, and both basic and advanced natural-language scenario planning.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (4.6 / 5), Lifesight (4.5 / 5), Mutinex (4.5 / 5), Triple Whale (3.6 / 5), Prescient AI (3.4 / 5), and Fospha (2.4 / 5). Budget-allocation recommendations, revenue forecasting, and basic scenario planning received the strongest average scores across the vendor set. Marginal incremental ROAS and advanced natural-language scenario planning received lower average scores, while the largest vendor-level differences appeared in budget allocation, revenue forecasting, and both forms of scenario planning.

Criteria-level scores for Category 3Each criterion is scored from 0 to 1; the category maximum is 5. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
3.1 AI recommends optimal budget allocation by channel 1.0 0.9 1.0 0.5 1.0 0.9 0.9
3.2 AI forecasts total revenue based on optimal budget allocation 1.0 0.8 0.9 0.5 1.0 0.8 0.8
3.3 AI provides miROAS and response curves for each channel 0.8 0.6 0.8 0.5 0.8 0.8 0.7
3.4 AI supports basic custom scenario planning (e.g., "what if I cut Meta by 20%?") 1.0 0.8 1.0 0.5 1.0 0.5 0.8
3.5 AI supports advanced scenario planning in natural language (constraints, multi-dimensional optimization…) 0.9 0.6 0.9 0.4 0.8 0.5 0.7
Category total out of 5 4.6 3.6 4.5 2.4 4.5 3.4 3.8

4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI

Category 4 research score by vendorCampaign & Ad Set-Level Optimization for Digital Channels with conversational AI (maximum 6)
Sellforte4.3 out of 64.3 / 6
Triple Whale3.5 out of 63.5 / 6
Lifesight2.9 out of 62.9 / 6
Fospha2.8 out of 62.8 / 6
Prescient AI2.1 out of 62.1 / 6
Mutinex1.1 out of 61.1 / 6

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category examines 6 criteria that assess whether the vendor's conversational AI can provide campaign- and ad set-level incremental and marginal ROAS, compare incrementality results with attribution metrics, recommend daily budgets and bids, and analyze results before and after executed changes.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (4.3 / 6), Triple Whale (3.5 / 6), Lifesight (2.9 / 6), Fospha (2.8 / 6), Prescient AI (2.1 / 6), and Mutinex (1.1 / 6). The vendors' scores were most closely grouped for campaign and ad set-level incremental ROAS. Greater separation appeared in optimal daily budget recommendations, optimal bid recommendations, and marginal ROAS at campaign and ad set level; the pre/post analysis criterion also received comparatively modest scores across the vendor set.

Criteria-level scores for Category 4Each criterion is scored from 0 to 1; the category maximum is 6. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
4.1 AI provides incremental ROAS of each campaign & ad set 0.8 0.5 0.6 0.8 0.8 0.8 0.7
4.2 AI compares incremental ROAS to last-click and ad platform attribution ROAS 0.6 0.8 0.6 0.4 0.1 0.5 0.5
4.3 AI provides miROAS for each campaign & ad set 0.8 0.3 0.5 0.3 0.1 0.4 0.4
4.4 AI recommends optimal spend for each campaign & ad set 0.9 0.8 0.6 0.6 0.1 0.4 0.6
4.5 AI recommends optimal bid value for each campaign & ad set 0.8 0.8 0.3 0.3 0.0 0.0 0.3
4.6 AI provides pre/post analysis for each bidding change 0.5 0.5 0.3 0.5 0.0 0.1 0.3
Category total out of 6 4.3 3.5 2.9 2.8 1.1 2.1 2.8

5. Incrementality Testing with conversational AI

Category 5 research score by vendorIncrementality Testing with conversational AI (maximum 5)
Sellforte3.8 out of 53.8 / 5
Triple Whale2.9 out of 52.9 / 5
Lifesight2.6 out of 52.6 / 5
Fospha1.4 out of 51.4 / 5
Mutinex0.6 out of 50.6 / 5
Prescient AI0.6 out of 50.6 / 5

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category examines 5 criteria that assess whether the vendor's conversational AI can summarize geo-test and own-media A/B test results, surface findings from Meta Conversion Lift tests, recommend what to test next, and provide incrementality test-design guidance.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (3.8 / 5), Triple Whale (2.9 / 5), Lifesight (2.6 / 5), Fospha (1.4 / 5), Mutinex (0.6 / 5), and Prescient AI (0.6 / 5). Geo-test reporting received the strongest average score across the six vendors. The own-media A/B testing and Meta Conversion Lift criteria produced larger differences, while recommendations about what to test and how to design a test sat between those two patterns.

Criteria-level scores for Category 5Each criterion is scored from 0 to 1; the category maximum is 5. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
5.1 AI can summarize results for Geo Tests 0.9 0.6 0.9 0.5 0.3 0.3 0.6
5.2 AI can summarize results for Own Media A/B tests (e.g. leaflet tests) 0.8 0.0 0.3 0.0 0.0 0.0 0.2
5.3 Platform ingests Meta Conversion Lift tests and AI summarizes findings 0.9 0.9 0.1 0.3 0.0 0.3 0.4
5.4 AI provides incrementality testing recommendations 0.6 0.9 0.8 0.3 0.3 0.1 0.5
5.5 AI provides incrementality test design recommendations 0.6 0.5 0.6 0.4 0.1 0.0 0.4
Category total out of 5 3.8 2.9 2.6 1.4 0.6 0.6 2.0

6. Agentic Execution & Autonomy

Category 6 research score by vendorAgentic Execution & Autonomy (maximum 4)
Triple Whale3.9 out of 43.9 / 4
Lifesight3.1 out of 43.1 / 4
Fospha2.4 out of 42.4 / 4
Sellforte2.1 out of 42.1 / 4
Prescient AI0.4 out of 40.4 / 4
Mutinex0.1 out of 40.1 / 4

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category examines 4 criteria that assess whether the vendor's conversational AI can initiate or execute budget and bidding changes through advertising-platform APIs, support configurable levels of autonomy, and proactively provide insights and alerts.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Triple Whale (3.9 / 4), Lifesight (3.1 / 4), Fospha (2.4 / 4), Sellforte (2.1 / 4), Prescient AI (0.4 / 4), and Mutinex (0.1 / 4). This category produced substantial differences on all four criteria: executing budget changes, executing bidding changes, configuring the level of autonomy, and providing proactive insights or alerts. The spread indicates that Claude and ChatGPT found materially different levels of public documentation across the vendor set for agentic execution and autonomy.

Criteria-level scores for Category 6Each criterion is scored from 0 to 1; the category maximum is 4. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
6.1 AI can push daily spend/budget changes to Meta, Google, TikTok APIs 0.8 1.0 0.9 0.8 0.0 0.0 0.6
6.2 AI can push bidding changes (e.g., Target ROAS) to Meta, Google, TikTok APIs 0.8 1.0 0.6 0.4 0.0 0.0 0.5
6.3 AI's level of autonomy for execution can be configured 0.6 0.9 0.9 0.6 0.0 0.0 0.5
6.4 Proactive insights & alerts via AI 0.0 1.0 0.8 0.6 0.1 0.4 0.5
Category total out of 4 2.1 3.9 3.1 2.4 0.1 0.4 2.0

7. AI's UX & Conversational Interface

Category 7 research score by vendorAI's UX & Conversational Interface (maximum 8)
Sellforte6.8 out of 86.8 / 8
Triple Whale6.8 out of 86.8 / 8
Fospha5.8 out of 85.8 / 8
Lifesight4.0 out of 84.0 / 8
Mutinex3.5 out of 83.5 / 8
Prescient AI3.5 out of 83.5 / 8

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category directly evaluates the conversational interface and related access surfaces across 8 criteria, including inline visualizations, conversation history, export options, links to deeper dashboards, embedded workflows, visible reasoning, non-English use, and access through MCP or an external LLM.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (6.8 / 8), Triple Whale (6.8 / 8), Fospha (5.8 / 8), Lifesight (4.0 / 8), Mutinex (3.5 / 8), and Prescient AI (3.5 / 8). Inline tables and charts, multi-turn context, export options, and MCP or external-LLM access received relatively strong scores across the vendor set. The largest differences appeared in visible reasoning, production use in non-English languages, MCP access, and links from AI outputs to deeper dashboards.

Criteria-level scores for Category 7Each criterion is scored from 0 to 1; the category maximum is 8. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
7.1 Includes tables and charts inline in AI responses 1.0 1.0 0.5 0.5 1.0 0.3 0.7
7.2 AI has conversation history & multi-turn context retention 1.0 1.0 0.3 1.0 1.0 0.8 0.8
7.3 Allows convenient export of AI outputs (e.g., PDF, Slides, CSV) 1.0 0.8 0.5 0.5 1.0 0.5 0.7
7.4 AI grounds answers in data, providing links from outputs to deep-dive dashboards 0.8 0.8 0.0 0.5 0.0 0.8 0.5
7.5 AI operates in embedded and multi-window mode (chat + dashboard side-by-side) 0.8 1.0 0.5 0.5 0.5 0.5 0.6
7.6 AI shows its reasoning steps and logic while answering 1.0 1.0 1.0 0.8 0.0 0.5 0.7
7.7 AI handles non-English questions in production 0.3 0.3 0.3 1.0 0.0 0.0 0.3
7.8 AI can be accessed via MCP server / external LLM 1.0 1.0 1.0 1.0 0.0 0.3 0.7
Category total out of 8 6.8 6.8 4.0 5.8 3.5 3.5 5.0

8. Analytical Backbone of the Conversational AI

Category 8 research score by vendorAnalytical Backbone of the Conversational AI (maximum 4)
Triple Whale3.8 out of 43.8 / 4
Prescient AI3.8 out of 43.8 / 4
Sellforte3.5 out of 43.5 / 4
Lifesight3.0 out of 43.0 / 4
Mutinex2.8 out of 42.8 / 4
Fospha2.5 out of 42.5 / 4

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category is an intentional exception to the conversational-availability focus of Categories 1–6: its 4 criteria evaluate the analytical backbone supporting the AI, rather than whether every underlying modeling capability is exposed through chat. The criteria cover model-backed answers, a Bayesian MMM foundation, calibration with incrementality tests, model-validation reporting, and auditable model configuration.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Triple Whale (3.8 / 4), Prescient AI (3.8 / 4), Sellforte (3.5 / 4), Lifesight (3.0 / 4), Mutinex (2.8 / 4), and Fospha (2.5 / 4). Calibration with incrementality tests and reporting of model-validation metrics received consistently strong scores across the vendor set. The largest differences came from documentation of the Bayesian or model-backed foundation and whether model configuration—such as priors—was auditable and editable in a self-serve interface.

Criteria-level scores for Category 8Each criterion is scored from 0 to 1; the category maximum is 4. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
8.1 AI provides deterministic, model-backed answers with Bayesian MMM as backbone 1.0 0.8 0.3 0.8 1.0 1.0 0.8
8.2 Bayesian MMM used by the AI is calibrated with incrementality tests 1.0 1.0 0.8 0.8 0.8 1.0 0.9
8.3 AI reports model validation and other modelling KPIs 1.0 1.0 1.0 0.8 0.8 1.0 0.9
8.4 Model calibration & configuration settings (e.g., priors) are auditable and editable in a self-serve UI 0.5 1.0 1.0 0.3 0.3 0.8 0.6
Category total out of 4 3.5 3.8 3.0 2.5 2.8 3.8 3.2

9. Enterprise-Grade Platform

Category 9 research score by vendorEnterprise-Grade Platform (maximum 7)
Sellforte6.3 out of 76.3 / 7
Lifesight5.0 out of 75.0 / 7
Mutinex3.5 out of 73.5 / 7
Fospha3.0 out of 73.0 / 7
Triple Whale2.0 out of 72.0 / 7
Prescient AI0.3 out of 70.3 / 7

Average of Claude and ChatGPT scores; public documentation reviewed as of September 2026.

What we evaluated. This category is an intentional exception to the conversational-availability focus of Categories 1–6: its 7 criteria evaluate the broader platform's enterprise readiness, rather than whether each capability is delivered through chat. The criteria cover enterprise references, independently audited security, regional data residency, multi-cloud options, SSO, controls on sharing customer data with third-party LLMs, and access to an ungated demo or trial.

What the research found. Within this evaluation, the scores by vendor, from highest to lowest, were Sellforte (6.3 / 7), Lifesight (5.0 / 7), Mutinex (3.5 / 7), Fospha (3.0 / 7), Triple Whale (2.0 / 7), and Prescient AI (0.3 / 7). Third-party security assurance and SSO were well documented for most, but not all, vendors. The greatest differences appeared in regional data residency, multi-cloud availability, controls on sending customer data to third-party LLMs, and access to an ungated hands-on demo or trial.

Criteria-level scores for Category 9Each criterion is scored from 0 to 1; the category maximum is 7. The Average column is the mean score across the six evaluated vendors. Scores reflect evidence Claude and ChatGPT identified in public materials reviewed as of September 2026, not direct product tests or confirmation that a vendor has or lacks a capability. Lower scores may reflect limited or unclear public documentation; verify current functionality directly with each vendor.
Evaluated criterion Sellforte Triple Whale Lifesight Fospha Mutinex Prescient AI Average
9.1 At least 10 public reference customers from $1B+ revenue brands 0.8 0.5 0.8 0.5 0.8 0.3 0.6
9.2 SOC 2, ISO 27001, or audited IT security by a third-party cyber security auditor 1.0 1.0 1.0 1.0 1.0 0.0 0.8
9.3 Data residency: geography option between US and EU 0.8 0.0 1.0 0.5 0.0 0.0 0.4
9.4 Multi-cloud: option between AWS, GCP, and Azure 1.0 0.0 0.0 0.0 0.0 0.0 0.2
9.5 Supports single sign-on (SSO) for enterprises 1.0 0.3 1.0 1.0 1.0 0.0 0.7
9.6 Customer data is not shared to a third-party LLM (LLM deployed in customer-specific cloud container) 0.8 0.0 0.8 0.0 0.8 0.0 0.4
9.7 Hands-on demo or trial of the AI is available without sales-call gating 1.0 0.3 0.5 0.0 0.0 0.0 0.3
Category total out of 7 6.3 2.0 5.0 3.0 3.5 0.3 3.3

What this evaluation suggests about the market

Average research score by categoryMean across six vendor scores, normalized to each category maximum (100%)
1. Marketing reporting89.6 percent89.6%
2. Historical insights86.3 percent86.3%
8. Analytical backbone80.2 percent80.2%
3. Channel optimization76.7 percent76.7%
7. AI UX63.0 percent63.0%
6. Agentic execution50.0 percent50.0%
9. Enterprise platform47.6 percent47.6%
4. Campaign optimization46.2 percent46.2%
5. Incrementality testing39.6 percent39.6%

Based on Claude and ChatGPT scoring of public documentation reviewed as of September 2026.

Across the six evaluated vendors, the strongest normalized category averages were marketing data reporting (89.6%), historical performance insights and causal explanation (86.3%), the analytical backbone (80.2%), and channel-level optimization (76.7%). In this research, that pattern suggests that public documentation was most consistent for turning measurement outputs into reporting, explanations, forecasting, and higher-level planning support. A strong category average does not mean that every product covered every workflow equally; it means Claude and ChatGPT found relevant evidence more consistently across the vendors and criteria included in that category.

The lowest normalized category averages were incrementality testing (39.6%), campaign and ad set-level optimization (46.2%), enterprise-grade platform capabilities (47.6%), and agentic execution and autonomy (50.0%). These results point to the greatest apparent improvement—or documentation—opportunity in workflows involving experimental design and integration, granular campaign action, execution through advertising platforms, and enterprise controls. Several of these categories also showed wider differences between vendors, so buyers should validate exact workflows, supported channels, governance controls, and current availability directly. Because the assessment covers six vendors and public sources rather than hands-on product testing, this chart is best read as a market signal from this evaluation, not a definitive measure of the industry.

Vendor commentary

1. Sellforte (research score: 39.8 out of 48)

Research overview

Sellforte received an average research score of 39.8 out of 48, the highest total in this six-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 3, 2026, across 48 criteria in nine categories.

Sellforte website screenshot

Category scorecard

Sellforte: scores assigned in this researchAverage and highest research scores cover all six vendors in this evaluation, including Sellforte. Calculated from each vendor's unrounded average of Claude and ChatGPT scores, then rounded to one decimal. The total row compares vendor totals. Scores show points received out of the available maximum.
CategorySellforteAverage research scoreHighest research score
1. Marketing Data Reporting via conversational AI3.5 / 43.6 / 43.9 / 4
2. Historical Performance Insights & Causal Explanation via conversational AI5.0 / 54.3 / 55.0 / 5
3. Channel-Level Optimization with conversational AI4.6 / 53.8 / 54.6 / 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI4.3 / 62.8 / 64.3 / 6
5. Incrementality Testing with conversational AI3.8 / 52.0 / 53.8 / 5
6. Agentic Execution & Autonomy2.1 / 42.0 / 43.9 / 4
7. AI's UX & Conversational Interface6.8 / 85.0 / 86.8 / 8
8. Analytical Backbone of the Conversational AI3.5 / 43.2 / 43.8 / 4
9. Enterprise-Grade Platform6.3 / 73.3 / 76.3 / 7
Total score out of 4839.8 / 4830.1 / 4839.8 / 48

Categories cover different numbers of criteria. The higher- and lower-scoring groups below describe Sellforte's own score profile as a share of available points. Comparisons with the six-vendor average describe its position within this research; a lower-scoring category in its own profile can still be above that average. See the scoring methodology.

Sellforte: category research scoresResearch score as a proportion of available points in each category
SellforteSix-vendor averageHighest research score
1. Marketing data reporting
2. Historical performance insights
3. Channel-level optimization
4. Campaign & ad set optimization
5. Incrementality testing
6. Agentic execution & autonomy
7. Conversational interface
8. Analytical backbone
9. Enterprise-grade platform

Percentage of available points

Bars show each vendor’s average Claude and ChatGPT research score as a percentage of the category maximum. Markers show the average and highest scores across all six vendors, including Sellforte. Positions use unrounded scores; labels show points rounded to one decimal.

Where Claude and ChatGPT assigned higher scores

  • Historical performance insights and causal explanation — 5.0 / 5. Its assigned score was above the six-vendor research average of 4.3 / 5. It received the highest assigned score in this category. In the criterion-level results, the criteria for incremental returns by digital channel, daily MMM-based measurement updates and explanations of performance changes each averaged 1 / 1.

  • Channel-level optimization — 4.6 / 5. Its assigned score was above the six-vendor research average of 3.8 / 5. It received the highest assigned score in this category. In the criterion-level results, the criteria for channel budget recommendations and revenue forecasts each averaged 1 / 1; the criterion for marginal returns and response curves averaged 0.75 / 1.

  • Enterprise-grade platform — 6.3 / 7. Its assigned score was above the six-vendor research average of 3.3 / 7. It received the highest assigned score in this category. In the criterion-level results, the criteria for security certification or third-party audits and single sign-on each averaged 1 / 1; the criterion for a choice of US or EU data residency averaged 0.75 / 1.

Where scores were lower or Claude and ChatGPT differed

  • Agentic execution and autonomy — 2.1 / 4. Its assigned score was above the six-vendor research average of 2.0 / 4. The highest recorded score was 3.9 / 4. In the criterion-level results, the criterion for executing budget changes averaged 0.75 / 1; the criterion for configurable autonomy averaged 0.625 / 1; the criterion for proactive insights and alerts averaged 0 / 1.

  • Campaign and ad set-level optimization — 4.3 / 6. Its assigned score was above the six-vendor research average of 2.8 / 6. It received the highest assigned score in this category. In the criterion-level results, the criterion for campaign and ad set daily-budget recommendations averaged 0.875 / 1; the criterion for analysis of revenue and spend before and after bidding changes averaged 0.5 / 1.

  • Incrementality testing — 3.8 / 5. Its assigned score was above the six-vendor research average of 2.0 / 5. It received the highest assigned score in this category. In the criterion-level results, the criterion for geographic-test summaries averaged 0.875 / 1; the criteria for test recommendations and test-design recommendations each averaged 0.625 / 1.

Different scores from Claude and ChatGPT: For agentic execution and autonomy, ChatGPT assigned 3 / 4 and Claude assigned 1.25 / 4. The research records show the separate criterion-level scores, rationales, and cited sources behind these assessments.

Research records and vendor response

See the evaluation research records for Sellforte's criterion-level scores, Claude and ChatGPT's separate rationales, and cited sources. Vendors can submit additional documentation for re-evaluation at research@sellforte.com.

Research summary

Sellforte received the highest total score in this comparison, at 39.8 / 48. It received the highest assigned scores in historical performance insights, channel-level optimization, campaign and ad set-level optimization, incrementality testing, and the enterprise category.

Agentic execution and autonomy was its lowest-scoring category as a proportion of available points, at 2.1 / 4, with different scores assigned by Claude and ChatGPT. These results summarize the scores assigned by Claude and ChatGPT under this research’s criteria and methodology.

2. Triple Whale (research score: 34.0 out of 48)

Research overview

Triple Whale received an average research score of 34.0 out of 48, the second-highest total in this six-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 3, 2026, across 48 criteria in nine categories.

Triple Whale website

Category scorecard

Triple Whale: scores assigned in this researchAverage and highest research scores cover all six vendors in this evaluation, including Sellforte. Calculated from each vendor's unrounded average of Claude and ChatGPT scores, then rounded to one decimal. The total row compares vendor totals. Scores show points received out of the available maximum.
Category Triple Whale Average research score Highest research score
1. Marketing Data Reporting via conversational AI 3.8 / 4 3.6 / 4 3.9 / 4
2. Historical Performance Insights & Causal Explanation via conversational AI 3.9 / 5 4.3 / 5 5.0 / 5
3. Channel-Level Optimization with conversational AI 3.6 / 5 3.8 / 5 4.6 / 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI 3.5 / 6 2.8 / 6 4.3 / 6
5. Incrementality Testing with conversational AI 2.9 / 5 2.0 / 5 3.8 / 5
6. Agentic Execution & Autonomy 3.9 / 4 2.0 / 4 3.9 / 4
7. AI's UX & Conversational Interface 6.8 / 8 5.0 / 8 6.8 / 8
8. Analytical Backbone of the Conversational AI 3.8 / 4 3.2 / 4 3.8 / 4
9. Enterprise-Grade Platform 2.0 / 7 3.3 / 7 6.3 / 7
Total score out of 48 34.0 / 48 30.1 / 48 39.8 / 48

Categories cover different numbers of criteria. The higher- and lower-scoring groups below describe Triple Whale's own score profile as a share of available points. Comparisons with the six-vendor average describe its position within this research; a lower-scoring category in its own profile can still be above that average. See the scoring methodology.

Triple Whale: category research scoresResearch score as a proportion of available points in each category
Triple WhaleSix-vendor averageHighest research score
1. Marketing data reporting
2. Historical performance insights
3. Channel-level optimization
4. Campaign & ad set optimization
5. Incrementality testing
6. Agentic execution & autonomy
7. Conversational interface
8. Analytical backbone
9. Enterprise-grade platform

Percentage of available points

Bars show each vendor’s average Claude and ChatGPT research score as a percentage of the category maximum. Markers show the average and highest scores across all six vendors, including Sellforte. Positions use unrounded scores; labels show points rounded to one decimal.

Where Claude and ChatGPT assigned higher scores

  • Agentic execution and autonomy — 3.9 / 4. Claude and ChatGPT each assigned 1 / 1 to the criteria covering budget changes, bidding changes, and proactive insights or alerts (6.1, 6.2, 6.4). Its assigned score was above the six-vendor research average of 2.0 / 4 and was the highest recorded score in this category.

  • Marketing data reporting — 3.8 / 4. Claude and ChatGPT each assigned 1 / 1 to online-sales reporting and digital-media reporting (1.1, 1.3). The offline-sales and offline-media criteria each averaged 0.875 / 1 (1.2, 1.4).

  • Analytical backbone — 3.8 / 4. Claude and ChatGPT each assigned 1 / 1 to the criteria for incrementality-test calibration, model-validation reporting, and auditable and editable model settings (8.2–8.4).

Where scores were lower or Claude and ChatGPT differed

  • Enterprise-grade platform — 2.0 / 7. This was Triple Whale's lowest category score as a share of available points, below the six-vendor research average of 3.3 / 7. The highest recorded score was 6.3 / 7. Within that total, Claude and ChatGPT each assigned 1 / 1 to the security certification or third-party audit criterion (9.2). The category total should be read alongside its individual criteria.

  • Incrementality testing — 2.9 / 5. Although this was one of Triple Whale's lower-scoring categories as a share of available points, its assigned score was above the six-vendor research average of 2.0 / 5 and below the highest recorded score of 3.8 / 5. The Meta Conversion Lift and test-recommendation criteria each averaged 0.875 / 1 (5.3, 5.4). Claude and ChatGPT each assigned 0 / 1 to the owned-media A/B-test summary criterion (5.2), recording no public evidence under the methodology.

  • Campaign and ad set-level optimization — 3.5 / 6. Its assigned score was above the six-vendor research average of 2.8 / 6 and below the highest recorded score of 4.3 / 6. Claude and ChatGPT each assigned 0.25 / 1 to the campaign and ad set-level marginal incremental ROAS criterion (4.3), and 0.75 / 1 to the daily-budget recommendation criterion (4.4).

  • Channel-level optimization — 3.6 / 5. ChatGPT assigned 4.5 / 5 and Claude assigned 2.75 / 5. The displayed average combines these different assessments; the research records contain the separate scores and rationales.

Research records and vendor response

See the evaluation research records for Triple Whale's criterion-level scores, Claude and ChatGPT's separate rationales, and cited sources. Vendors can submit additional documentation for re-evaluation at research@sellforte.com.

Research summary

Triple Whale received the second-highest total score in this comparison, at 34.0 / 48. Its assigned score for agentic execution and autonomy was the highest among the six vendors, and it tied for the highest scores in conversational interface and analytical backbone.

Its incrementality-testing and campaign-level optimization scores were above the six-vendor research averages, while its enterprise-category score was below the average. These results summarize the scores assigned by Claude and ChatGPT under this research’s criteria and methodology.

3. Lifesight (research score: 32.9 out of 48)

Research overview

Lifesight received an average research score of 32.9 out of 48, the third-highest total in this six-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 3, 2026, across 48 criteria in nine categories.

Lifesight website

Category scorecard

Lifesight: scores assigned in this researchAverage and highest research scores cover all six vendors in this evaluation, including Sellforte. Calculated from each vendor's unrounded average of Claude and ChatGPT scores, then rounded to one decimal. The total row compares vendor totals. Scores show points received out of the available maximum.
CategoryLifesightAverage research scoreHighest research score
1. Marketing Data Reporting via conversational AI3.9 / 43.6 / 43.9 / 4
2. Historical Performance Insights & Causal Explanation via conversational AI3.9 / 54.3 / 55.0 / 5
3. Channel-Level Optimization with conversational AI4.5 / 53.8 / 54.6 / 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI2.9 / 62.8 / 64.3 / 6
5. Incrementality Testing with conversational AI2.6 / 52.0 / 53.8 / 5
6. Agentic Execution & Autonomy3.1 / 42.0 / 43.9 / 4
7. AI's UX & Conversational Interface4.0 / 85.0 / 86.8 / 8
8. Analytical Backbone of the Conversational AI3.0 / 43.2 / 43.8 / 4
9. Enterprise-Grade Platform5.0 / 73.3 / 76.3 / 7
Total score out of 4832.9 / 4830.1 / 4839.8 / 48

Categories cover different numbers of criteria. The higher- and lower-scoring groups below describe Lifesight's own score profile as a share of available points. Comparisons with the six-vendor average describe its position within this research; a lower-scoring category in its own profile can still be above that average. See the scoring methodology.

Lifesight: category research scoresResearch score as a proportion of available points in each category
LifesightSix-vendor averageHighest research score
1. Marketing data reporting
2. Historical performance insights
3. Channel-level optimization
4. Campaign & ad set optimization
5. Incrementality testing
6. Agentic execution & autonomy
7. Conversational interface
8. Analytical backbone
9. Enterprise-grade platform

Percentage of available points

Bars show each vendor’s average Claude and ChatGPT research score as a percentage of the category maximum. Markers show the average and highest scores across all six vendors, including Sellforte. Positions use unrounded scores; labels show points rounded to one decimal.

Where Claude and ChatGPT assigned higher scores

  • Marketing data reporting — 3.9 / 4. Its assigned score was above the six-vendor research average of 3.6 / 4. It tied for the highest assigned score in this category. In the criterion-level results, the criteria for online-sales reporting and offline-sales reporting each averaged 1 / 1; the criterion for offline-media reporting averaged 0.875 / 1.

  • Channel-level optimization — 4.5 / 5. Its assigned score was above the six-vendor research average of 3.8 / 5. The highest recorded score was 4.6 / 5. In the criterion-level results, the criteria for channel budget recommendations and basic scenario planning each averaged 1 / 1; the criterion for advanced scenario planning averaged 0.875 / 1.

  • Agentic execution and autonomy — 3.1 / 4. Its assigned score was above the six-vendor research average of 2.0 / 4. The highest recorded score was 3.9 / 4. In the criterion-level results, the criteria for executing budget changes and configurable autonomy each averaged 0.875 / 1; the criterion for executing bidding changes averaged 0.625 / 1.

Where scores were lower or Claude and ChatGPT differed

  • Campaign and ad set-level optimization — 2.9 / 6. Its assigned score was above the six-vendor research average of 2.8 / 6. The highest recorded score was 4.3 / 6. In the criterion-level results, the criterion for campaign and ad set incremental ROAS averaged 0.625 / 1; the criteria for bidding-target recommendations and analysis before and after bidding changes each averaged 0.25 / 1.

  • Conversational interface and user experience — 4.0 / 8. Its assigned score was below the six-vendor research average of 5.0 / 8. The highest recorded score was 6.8 / 8. In the criterion-level results, the criterion for showing the reasoning behind answers averaged 1 / 1; the criterion for conversation history and multi-turn context averaged 0.25 / 1; the criterion for links from answers to detailed dashboards averaged 0 / 1.

  • Incrementality testing — 2.6 / 5. Its assigned score was above the six-vendor research average of 2.0 / 5. The highest recorded score was 3.8 / 5. In the criterion-level results, the criterion for geographic-test summaries averaged 0.875 / 1; the criterion for Meta Conversion Lift ingestion and AI summaries averaged 0.125 / 1; the criterion for owned-media A/B-test summaries averaged 0.25 / 1.

Different scores from Claude and ChatGPT: For analytical backbone, ChatGPT assigned 2.5 / 4 and Claude assigned 3.5 / 4. The research records show the separate criterion-level scores, rationales, and cited sources behind these assessments.

Research records and vendor response

See the evaluation research records for Lifesight's criterion-level scores, Claude and ChatGPT's separate rationales, and cited sources. Vendors can submit additional documentation for re-evaluation at research@sellforte.com.

Research summary

Lifesight received the third-highest total score in this comparison, at 32.9 / 48. It tied for the highest assigned score in marketing data reporting and the second-highest score in channel-level optimization.

Its incrementality-testing and enterprise-category scores were above the six-vendor research averages, while its conversational-interface score was below the average. These results summarize the scores assigned by Claude and ChatGPT under this research’s criteria and methodology.

4. Fospha (research score: 26.9 out of 48)

Research overview

Fospha received an average research score of 26.9 out of 48, the fourth-highest total in this six-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 3, 2026, across 48 criteria in nine categories.

Fospha website

Category scorecard

Fospha: scores assigned in this researchAverage and highest research scores cover all six vendors in this evaluation, including Sellforte. Calculated from each vendor's unrounded average of Claude and ChatGPT scores, then rounded to one decimal. The total row compares vendor totals. Scores show points received out of the available maximum.
CategoryFosphaAverage research scoreHighest research score
1. Marketing Data Reporting via conversational AI2.6 / 43.6 / 43.9 / 4
2. Historical Performance Insights & Causal Explanation via conversational AI4.1 / 54.3 / 55.0 / 5
3. Channel-Level Optimization with conversational AI2.4 / 53.8 / 54.6 / 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI2.8 / 62.8 / 64.3 / 6
5. Incrementality Testing with conversational AI1.4 / 52.0 / 53.8 / 5
6. Agentic Execution & Autonomy2.4 / 42.0 / 43.9 / 4
7. AI's UX & Conversational Interface5.8 / 85.0 / 86.8 / 8
8. Analytical Backbone of the Conversational AI2.5 / 43.2 / 43.8 / 4
9. Enterprise-Grade Platform3.0 / 73.3 / 76.3 / 7
Total score out of 4826.9 / 4830.1 / 4839.8 / 48

Categories cover different numbers of criteria. The higher- and lower-scoring groups below describe Fospha's own score profile as a share of available points. Comparisons with the six-vendor average describe its position within this research; a lower-scoring category in its own profile can still be above that average. See the scoring methodology.

Fospha: category research scoresResearch score as a proportion of available points in each category
FosphaSix-vendor averageHighest research score
1. Marketing data reporting
2. Historical performance insights
3. Channel-level optimization
4. Campaign & ad set optimization
5. Incrementality testing
6. Agentic execution & autonomy
7. Conversational interface
8. Analytical backbone
9. Enterprise-grade platform

Percentage of available points

Bars show each vendor’s average Claude and ChatGPT research score as a percentage of the category maximum. Markers show the average and highest scores across all six vendors, including Sellforte. Positions use unrounded scores; labels show points rounded to one decimal.

Where Claude and ChatGPT assigned higher scores

  • Historical performance insights and causal explanation — 4.1 / 5. Its assigned score was below the six-vendor research average of 4.3 / 5. The highest recorded score was 5.0 / 5. In the criterion-level results, the criteria for incremental returns by digital channel and daily MMM-based measurement updates each averaged 1 / 1; the criterion for promotion-driven revenue reporting averaged 0.25 / 1.

  • Conversational interface and user experience — 5.8 / 8. Its assigned score was above the six-vendor research average of 5.0 / 8. The highest recorded score was 6.8 / 8. In the criterion-level results, the criteria for conversation history and multi-turn context and non-English questions each averaged 1 / 1; the criterion for tables and charts inside answers averaged 0.5 / 1.

  • Marketing data reporting — 2.6 / 4. Its assigned score was below the six-vendor research average of 3.6 / 4. The highest recorded score was 3.9 / 4. In the criterion-level results, the criteria for online-sales reporting and digital-media reporting each averaged 1 / 1; the criterion for offline-media reporting averaged 0.25 / 1.

Where scores were lower or Claude and ChatGPT differed

  • Incrementality testing — 1.4 / 5. Its assigned score was below the six-vendor research average of 2.0 / 5. The highest recorded score was 3.8 / 5. In the criterion-level results, the criterion for geographic-test summaries averaged 0.5 / 1; the criterion for Meta Conversion Lift ingestion and AI summaries averaged 0.25 / 1; the criterion for owned-media A/B-test summaries averaged 0 / 1.

  • Enterprise-grade platform — 3.0 / 7. Its assigned score was below the six-vendor research average of 3.3 / 7. The highest recorded score was 6.3 / 7. In the criterion-level results, the criteria for security certification or third-party audits and single sign-on each averaged 1 / 1; the criterion for a choice of US or EU data residency averaged 0.5 / 1.

  • Campaign and ad set-level optimization — 2.8 / 6. Its assigned score was below the six-vendor research average of 2.8 / 6. The highest recorded score was 4.3 / 6. In the criterion-level results, the criterion for campaign and ad set incremental ROAS averaged 0.75 / 1; the criterion for daily-budget recommendations averaged 0.625 / 1; the criterion for bidding-target recommendations averaged 0.25 / 1.

Different scores from Claude and ChatGPT: For analytical backbone, ChatGPT assigned 3.5 / 4 and Claude assigned 1.5 / 4. The research records show the separate criterion-level scores, rationales, and cited sources behind these assessments.

Research records and vendor response

See the evaluation research records for Fospha's criterion-level scores, Claude and ChatGPT's separate rationales, and cited sources. Vendors can submit additional documentation for re-evaluation at research@sellforte.com.

Research summary

Fospha received the fourth-highest total score in this comparison, at 26.9 / 48. Historical performance insights and causal explanation was its highest-scoring category as a proportion of available points, although its assigned score was below the six-vendor research average for that category.

Its conversational-interface and agentic-execution scores were above the six-vendor research averages. These results summarize the scores assigned by Claude and ChatGPT under this research’s criteria and methodology.

5. Mutinex (research score: 24.3 out of 48)

Research overview

Mutinex received an average research score of 24.3 out of 48, the fifth-highest total in this six-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 3, 2026, across 48 criteria in nine categories.

Mutinex website

Category scorecard

Mutinex: scores assigned in this researchAverage and highest research scores cover all six vendors in this evaluation, including Sellforte. Calculated from each vendor's unrounded average of Claude and ChatGPT scores, then rounded to one decimal. The total row compares vendor totals. Scores show points received out of the available maximum.
CategoryMutinexAverage research scoreHighest research score
1. Marketing Data Reporting via conversational AI3.9 / 43.6 / 43.9 / 4
2. Historical Performance Insights & Causal Explanation via conversational AI4.3 / 54.3 / 55.0 / 5
3. Channel-Level Optimization with conversational AI4.5 / 53.8 / 54.6 / 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI1.1 / 62.8 / 64.3 / 6
5. Incrementality Testing with conversational AI0.6 / 52.0 / 53.8 / 5
6. Agentic Execution & Autonomy0.1 / 42.0 / 43.9 / 4
7. AI's UX & Conversational Interface3.5 / 85.0 / 86.8 / 8
8. Analytical Backbone of the Conversational AI2.8 / 43.2 / 43.8 / 4
9. Enterprise-Grade Platform3.5 / 73.3 / 76.3 / 7
Total score out of 4824.3 / 4830.1 / 4839.8 / 48

Categories cover different numbers of criteria. The higher- and lower-scoring groups below describe Mutinex's own score profile as a share of available points. Comparisons with the six-vendor average describe its position within this research; a lower-scoring category in its own profile can still be above that average. See the scoring methodology.

Mutinex: category research scoresResearch score as a proportion of available points in each category
MutinexSix-vendor averageHighest research score
1. Marketing data reporting
2. Historical performance insights
3. Channel-level optimization
4. Campaign & ad set optimization
5. Incrementality testing
6. Agentic execution & autonomy
7. Conversational interface
8. Analytical backbone
9. Enterprise-grade platform

Percentage of available points

Bars show each vendor’s average Claude and ChatGPT research score as a percentage of the category maximum. Markers show the average and highest scores across all six vendors, including Sellforte. Positions use unrounded scores; labels show points rounded to one decimal.

Where Claude and ChatGPT assigned higher scores

  • Marketing data reporting — 3.9 / 4. Its assigned score was above the six-vendor research average of 3.6 / 4. It tied for the highest assigned score in this category. In the criterion-level results, the criteria for online-sales reporting, digital-media reporting and offline-media reporting each averaged 1 / 1.

  • Channel-level optimization — 4.5 / 5. Its assigned score was above the six-vendor research average of 3.8 / 5. The highest recorded score was 4.6 / 5. In the criterion-level results, the criteria for channel budget recommendations and revenue forecasts each averaged 1 / 1; the criterion for advanced scenario planning averaged 0.75 / 1.

  • Historical performance insights and causal explanation — 4.3 / 5. Its assigned score was below the six-vendor research average of 4.3 / 5. The highest recorded score was 5.0 / 5. In the criterion-level results, the criteria for incremental returns by digital channel and promotion-driven revenue reporting each averaged 1 / 1; the criterion for daily MMM-based measurement updates averaged 0.25 / 1.

Where scores were lower or Claude and ChatGPT differed

  • Agentic execution and autonomy — 0.1 / 4. Its assigned score was below the six-vendor research average of 2.0 / 4. The highest recorded score was 3.9 / 4. In the criterion-level results, the criteria for executing budget changes and executing bidding changes each averaged 0 / 1; the criterion for proactive insights and alerts averaged 0.125 / 1.

  • Incrementality testing — 0.6 / 5. Its assigned score was below the six-vendor research average of 2.0 / 5. The highest recorded score was 3.8 / 5. In the criterion-level results, the criteria for geographic-test summaries and test recommendations each averaged 0.25 / 1; the criterion for Meta Conversion Lift ingestion and AI summaries averaged 0 / 1.

  • Campaign and ad set-level optimization — 1.1 / 6. Its assigned score was below the six-vendor research average of 2.8 / 6. The highest recorded score was 4.3 / 6. In the criterion-level results, the criterion for campaign and ad set incremental ROAS averaged 0.75 / 1; the criterion for daily-budget recommendations averaged 0.125 / 1; the criterion for bidding-target recommendations averaged 0 / 1.

Different scores from Claude and ChatGPT: For campaign and ad set-level optimization, ChatGPT assigned 1.5 / 6 and Claude assigned 0.75 / 6. The research records show the separate criterion-level scores, rationales, and cited sources behind these assessments.

Research records and vendor response

See the evaluation research records for Mutinex's criterion-level scores, Claude and ChatGPT's separate rationales, and cited sources. Vendors can submit additional documentation for re-evaluation at research@sellforte.com.

Research summary

Mutinex received the fifth-highest total score in this comparison, at 24.3 / 48. It tied for the highest assigned score in marketing data reporting and the second-highest score in channel-level optimization.

Its campaign and ad set-level optimization, incrementality-testing, and agentic-execution scores were below the six-vendor research averages. These results summarize the scores assigned by Claude and ChatGPT under this research’s criteria and methodology.

6. Prescient AI (research score: 22.6 out of 48)

Research overview

Prescient AI received an average research score of 22.6 out of 48, the sixth-highest total in this six-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 4, 2026, across 48 criteria in nine categories.

Prescient AI website screenshot 2026-09-04

Category scorecard

Prescient AI: scores assigned in this researchAverage and highest research scores cover all six vendors in this evaluation, including Sellforte. Calculated from each vendor's unrounded average of Claude and ChatGPT scores, then rounded to one decimal. The total row compares vendor totals. Scores show points received out of the available maximum.
CategoryPrescient AIAverage research scoreHighest research score
1. Marketing Data Reporting via conversational AI3.9 / 43.6 / 43.9 / 4
2. Historical Performance Insights & Causal Explanation via conversational AI4.8 / 54.3 / 55.0 / 5
3. Channel-Level Optimization with conversational AI3.4 / 53.8 / 54.6 / 5
4. Campaign & Ad Set-Level Optimization for Digital Channels with conversational AI2.1 / 62.8 / 64.3 / 6
5. Incrementality Testing with conversational AI0.6 / 52.0 / 53.8 / 5
6. Agentic Execution & Autonomy0.4 / 42.0 / 43.9 / 4
7. AI's UX & Conversational Interface3.5 / 85.0 / 86.8 / 8
8. Analytical Backbone of the Conversational AI3.8 / 43.2 / 43.8 / 4
9. Enterprise-Grade Platform0.3 / 73.3 / 76.3 / 7
Total score out of 4822.6 / 4830.1 / 4839.8 / 48

Categories cover different numbers of criteria. The higher- and lower-scoring groups below describe Prescient AI's own score profile as a share of available points. Comparisons with the six-vendor average describe its position within this research; a lower-scoring category in its own profile can still be above that average. See the scoring methodology.

Prescient AI: category research scoresResearch score as a proportion of available points in each category
Prescient AISix-vendor averageHighest research score
1. Marketing data reporting
2. Historical performance insights
3. Channel-level optimization
4. Campaign & ad set optimization
5. Incrementality testing
6. Agentic execution & autonomy
7. Conversational interface
8. Analytical backbone
9. Enterprise-grade platform

Percentage of available points

Bars show each vendor’s average Claude and ChatGPT research score as a percentage of the category maximum. Markers show the average and highest scores across all six vendors, including Sellforte. Positions use unrounded scores; labels show points rounded to one decimal.

Where Claude and ChatGPT assigned higher scores

  • Marketing data reporting — 3.9 / 4. Its assigned score was above the six-vendor research average of 3.6 / 4. It tied for the highest assigned score in this category. In the criterion-level results, the criteria for online-sales reporting and offline-sales reporting each averaged 1 / 1; the criterion for offline-media reporting averaged 0.875 / 1.

  • Historical performance insights and causal explanation — 4.8 / 5. Its assigned score was above the six-vendor research average of 4.3 / 5. The highest recorded score was 5.0 / 5. In the criterion-level results, the criteria for incremental returns by digital channel and promotion-driven revenue reporting each averaged 1 / 1; the criterion for daily MMM-based measurement updates averaged 0.75 / 1.

  • Analytical backbone — 3.8 / 4. Its assigned score was above the six-vendor research average of 3.2 / 4. It tied for the highest assigned score in this category. In the criterion-level results, the criteria for model-backed answers with Bayesian MMM and calibration with incrementality tests each averaged 1 / 1; the criterion for auditable and editable model settings averaged 0.75 / 1.

Where scores were lower or Claude and ChatGPT differed

  • Enterprise-grade platform — 0.3 / 7. Its assigned score was below the six-vendor research average of 3.3 / 7. The highest recorded score was 6.3 / 7. In the criterion-level results, the criterion for public references from at least ten brands with revenue above $1 billion averaged 0.25 / 1; the criteria for security certification or third-party audits and single sign-on each averaged 0 / 1.

  • Agentic execution and autonomy — 0.4 / 4. Its assigned score was below the six-vendor research average of 2.0 / 4. The highest recorded score was 3.9 / 4. In the criterion-level results, the criteria for executing budget changes and executing bidding changes each averaged 0 / 1; the criterion for proactive insights and alerts averaged 0.375 / 1.

  • Incrementality testing — 0.6 / 5. Its assigned score was below the six-vendor research average of 2.0 / 5. The highest recorded score was 3.8 / 5. In the criterion-level results, the criteria for geographic-test summaries and Meta Conversion Lift ingestion and AI summaries each averaged 0.25 / 1; the criterion for test-design recommendations averaged 0 / 1.

Different scores from Claude and ChatGPT: For campaign and ad set-level optimization, ChatGPT assigned 1.5 / 6 and Claude assigned 2.75 / 6. The research records show the separate criterion-level scores, rationales, and cited sources behind these assessments.

Research records and vendor response

See the evaluation research records for Prescient AI's criterion-level scores, Claude and ChatGPT's separate rationales, and cited sources. Vendors can submit additional documentation for re-evaluation at research@sellforte.com.

Research summary

Prescient AI received the sixth-highest total score in this comparison, at 22.6 / 48. It received the second-highest assigned score in historical performance insights and causal explanation, and tied for the highest scores in marketing data reporting and analytical backbone.

Its enterprise-category, agentic-execution, and incrementality-testing scores were below the six-vendor research averages. These results summarize the scores assigned by Claude and ChatGPT under this research’s criteria and methodology.

Frequently Asked Questions (FAQ)

1. What are Conversational AI tools for MMM and Incrementality Testing?

AI tools for MMM and Incrementality Testing help marketers measure media performance, optimize marketing spend allocation and execute optimization actions through a conversational, natural language interface.

2. What are the best-performing Conversational AI tools for MMM and Incrementality Testing?

Sellforte, Triple Whale and Lifesight are currently the leading AI tools for MMM and Incrementality Testing, based on the 48-criteria evaluation in this article.

3. Which Conversational AI tool for MMM and Incrementality Testing is best for enterprise brands?

Based on this evaluation's LLM-based scoring, Sellforte received the highest scores for enterprise brands. Claude and ChatGPT found that Sellforte's conversational AI covers enterprise use-cases from channel-level spend optimization to tactical campaign-level bidding parameter optimization. Sellforte scored 6.3 out of 7 on the Enterprise-Grade Platform category, with Claude and ChatGPT finding public evidence of a large list of reference customers from $1B+ revenue brands, high-grade IT security, US and EU data residency options, multi-cloud deployment across AWS, GCP, and Azure, SSO, and LLM inference isolated inside customer-specific cloud infrastructure.

4. What is the difference between a conversational AI tool for MMM and Incrementality Testing and a generic AI chatbot like ChatGPT or Claude?

Generic AI chatbots like ChatGPT, Claude, and Gemini are not connected to your actual marketing data or measurement models. They can answer general questions about marketing concepts, but they cannot tell you what your incremental ROAS was on Meta last month, recommend an optimal budget allocation for your specific channels, or execute a bidding change on Google Ads. Conversational AI tools for MMM and Incrementality Testing are purpose-built for marketers, grounded in the advertiser's actual data and connected to incrementality-based measurement models such as Bayesian MMM and geo lift experiments.

5. Which conversational AI tool is best for campaign and ad set-level optimization?

Sellforte received the highest scores in this comparison for campaign and ad set-level optimization capabilities. Based on publicly available documentation, Claude and ChatGPT found evidence that Sellforte surfaces incremental ROAS for each campaign and ad set, compares it to last-click and platform-reported ROAS, and recommends optimal spend and bidding parameters at the campaign and ad set level.

6. Which conversational AI tool for MMM and Incrementality Testing is best for ecommerce and DTC brands?

The best choice depends on the size and requirements of the ecommerce brand. Sellforte is a strong option for mid-market and enterprise ecommerce brands (roughly $50M+ in revenue) that need both channel- and campaign-level optimization grounded in enterprise-grade MMM. Triple Whale is a strong option for small and mid-sized DTC brands that prioritize a strong AI and conversational UX but do not yet require enterprise-grade MMM or campaign-level incrementality optimization.

7. What is agentic MMM, and which tools support it?

Agentic MMM refers to AI tools that have a robust Marketing Mix Modeling system under the hood, which is accessed by specialized execution-focused agents. These agents recommend media spend allocation changes, bidding changes, and can autonomously execute those changes on ad platforms. Strongest Agentic MMM platforms in this research were Sellforte and Triple Whale. 

8. Sellforte vs. Triple Whale: Which is a better conversational AI tool for MMM and Incrementality Testing?

Sellforte and Triple Whale are the two highest-scoring tools in this comparison, but they serve different needs. Sellforte scored 39.8 out of 48, and Triple Whale  scored 34.0 out of 48.

According to this evaluation's LLM-based scoring, Sellforte leads in robustness of analytics, with Claude and ChatGPT finding strong public evidence of full-scale enterprise-grade MMM, incrementality testing and incrementality-corrected attribution. This enables Sellforte provide high-quality spend optimization on campaign & ad set level through its conversational AI interface. Triple Whale leads in agentic execution & autonomy, as well as conversational UX. Triple Whale is the better fit for smaller ecommerce brands that prioritize agentic execution and a strong conversational UX, but who do not require best-in-class incrementality based analytics. Sellforte is the better fit for mid-market and enterprise-grade marketing teams requiring best-in-class analytics and robust spend optimization.

9. What is Bayesian MMM, and why does it matter for AI tools in marketing measurement?

Bayesian Marketing Mix Modeling (MMM) is a statistical methodology for measuring the causal impact of marketing spend on sales. Unlike last-click attribution or platform-reported ROAS, which systematically overstate the contribution of lower-funnel channels, Bayesian MMM estimates true incremental impact across all channels simultaneously. For AI tools, Bayesian MMM is the analytical backbone that enables credible optimization recommendations. Without it, an AI can only report raw data rather than recommend where to allocate budget to maximize incremental revenue.

10. How were the 48 evaluation criteria in this comparison derived?

The 48 criteria were derived from three sources. First, 1,660 real prompts made by marketers and marketing analytics professionals in Sellforte, categorized by topic and use case. Second, AI-related statements and requirements from more than 700 discussions with Sellforte customers and prospects, including marketers, analytics leads, and data scientists from retail, ecommerce, DTC, travel, and restaurants — as well as recent marketing measurement RFP documentation. Third, a review of 78 vendors in the marketing measurement space to identify the range of existing AI capabilities. The combination of these three lenses produced a framework grounded in real usage patterns and real buyer expectations, rather than vendor marketing claims.

11. What is the difference between channel-level optimization and campaign and ad set-level optimization in AI tools?

Channel-level optimization means that the AI recommends how to allocate media budget across channels, for example, shifting spend from Meta to YouTube. This is where most AI tools in this comparison operate. Campaign and ad set-level optimization goes one level deeper: the AI recommends optimal spend and bidding parameters for each individual campaign and ad set within a channel, for example, recommending a specific Target ROAS value for each Google Ads campaign. Campaign and ad set-level optimization is more practically actionable because it is the level at which budget changes are actually executed in advertising platforms. It is also technically harder, requiring reliable incremental ROAS estimation at a much more granular level.

Change log

2026 May 26. Research was launched

2026 June 10. Research methodology was changed from Sellforte-based scoring to LLM-based scoring to minimize author bias. Evaluation for vendors was updated.

2026 Sep 3. Scoring was updated using the latest frontier LLM models: ChatGPT 5.6 Sol with Extra High effort, and Claude Fable 5.1 with High effort.

2026 Sep 4. Prescient AI was added to the evaluation.

Evaluation Dates by Vendor and LLM

The table below shows the exact dates when each vendor was evaluated by each LLM. 

Evaluation dates for Conversational AI Tools for MMM and Incrementality Testing Dates reflect when each LLM completed its scoring for the vendor
Vendor LLM conducting the evaluation Evaluation Date
Fospha ChatGPT 5.6 Sol ExtraHigh 3 September 2026
Fospha Claude Fable 5.1 High 3 September 2026
Lifesight ChatGPT 5.6 Sol ExtraHigh 3 September 2026
Lifesight Claude Fable 5.1 High 3 September 2026
Mutinex ChatGPT 5.6 Sol ExtraHigh 3 September 2026
Mutinex Claude Fable 5.1 High 3 September 2026
Sellforte ChatGPT 5.6 Sol ExtraHigh 3 September 2026
Sellforte Claude Fable 5.1 High 3 September 2026
Triple Whale ChatGPT 5.6 Sol ExtraHigh 3 September 2026
Triple Whale Claude Fable 5.1 High 3 September 2026
Prescient AI ChatGPT 5.6 Sol ExtraHigh 4 September 2026
Prescient AI Claude Fable 5.1 High 4 September 2026

Limitations & Disclosures

Author affiliation. This comparison is published by Sellforte, one of the platforms evaluated. We have made every effort to design the evaluation methodology fairly, applying the same evidence standards and scoring instructions to all vendors including ourselves. The actual scoring was conducted by Claude and ChatGPT, not by Sellforte. We encourage cross-referencing our research with vendor documentation, customer references, and independent analyst coverage.

Scope. This comparison primarily focuses on recognized commercial vendors. 

Snapshot in time. Vendor capabilities evolve quickly. This comparison reflects publicly available information as of dates specified in the Evaluation dates -section. Features released after that date are not yet reflected.

Corrections welcome. We will revise this comparison regularly. If a vendor believes a score is inaccurate, we welcome corrections with supporting documentation. Please email research@sellforte.com

Recommendation for buyers. Use this comparison as a structured starting point for your own evaluation, not as a final answer. We strongly recommend conducting evaluation calls or demos with vendors to verify fit against your organization's specific requirements. Pricing, services model, regional support, and integrations with your data infrastructure are factors no rubric can fully capture.

Further Reading & Resources

Methodology

AI and Agents in Marketing Measurement

Use-cases on media spend optimization

Original Marketing Measurement Research by Sellforte Labs

Practical hands-on guides

Marketing Measurement Tools, software and vendors

MMM tools, software and vendors for Ecommerce

MMM tools, software and vendors more broadly

Incrementality testing tools

Sellforte Product Features

Playbooks and Research from the industry

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte, with over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world’s largest advertisers on data-driven marketing optimization. Follow Lauri in LinkedIn, where he is one of the leading voices in MMM and marketing measurement.

Emil Kauppi-Hoyer

Emil Kauppi-Hoyer is Sellforte's Lead AI Engineer, leading the development of Sellforte AI. With a data science background, Emil belongs to Sellforte's engineering leadership. During his Sellforte career, Emil has implemented Marketing Mix Models and incrementality testing solutions to Sellforte's customers, while at the same time developing Sellforte's AI capabilities. Follow Emil in LinkedIn.

Juha Nuutinen

Juha Nuutinen is the Chief Executive Officer and co-founder at Sellforte, with over 15 years of experience in optimizing marketing spend and promotional activity for the largest advertisers in the world. Before co-founding Sellforte, he worked as a management consultant at the Boston Consulting Group, specializing in promotion optimization. Follow Juha in LinkedIn, where he is actively sharing his views on marketing measurement.