Top MCP Servers for Marketing Mix Modeling (MMM) and Incrementality Testing (2026 Q3)
Research disclosure: Sellforte designed and published this evaluation and is one of the vendors assessed. Claude and ChatGPT assigned the scores using public sources reviewed on September 23, 2026; this was not a hands-on product test. A score of zero means no public evidence was found under the methodology, not that a capability is absent. Vendors can submit documentation for re-evaluation at research@sellforte.com.
Quick Summary: Top MCP Servers for MMM and Incrementality Testing
This research compares eight vendors against 48 criteria across eight categories.
Sellforte received the highest average research score from Claude and ChatGPT at 38.3 out of 48, followed by Recast at 31.4 and Triple Whale at 30.6.
Research scores by vendor and category
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Blue shading indicates the share of available points; green outlines mark the highest scores, including ties.
| Research category | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam |
|---|---|---|---|---|---|---|---|---|
| Total research score / 48 | 38.3 | 31.4 | 30.6 | 29.8 | 26.6 | 24.3 | 18.9 | 17.6 |
| 1. Marketing Data Reporting via MCP / 6 | 6.0 | 5.0 | 5.3 | 4.1 | 4.9 | 4.3 | 3.1 | 5.0 |
| 2. Historical Performance Insights & Causal Explanation via MCP / 5 | 5.0 | 4.1 | 3.6 | 3.8 | 3.6 | 2.4 | 1.9 | 1.6 |
| 3. Channel-Level Optimization with MCP / 8 | 7.4 | 7.9 | 6.5 | 5.6 | 4.5 | 4.6 | 2.9 | 2.3 |
| 4. Campaign & Ad Set-Level Optimization with MCP / 8 | 5.0 | 1.0 | 3.8 | 4.1 | 2.4 | 3.3 | 3.1 | 1.4 |
| 5. Incrementality Testing with MCP / 5 | 2.6 | 3.6 | 2.5 | 2.1 | 3.0 | 1.3 | 2.6 | 1.1 |
| 6. MCP interoperability and workflows / 5 | 4.3 | 4.0 | 3.8 | 3.5 | 3.0 | 3.8 | 1.3 | 3.5 |
| 7. Analytical Backbone / 5 | 3.5 | 3.8 | 3.8 | 3.5 | 3.3 | 2.0 | 2.3 | 1.5 |
| 8. Enterprise-Grade Platform / 6 | 4.5 | 2.0 | 1.5 | 3.0 | 2.0 | 2.8 | 1.8 | 1.3 |
Research takeaways by vendor
- Sellforte: 38.3/48. Highest total score and the highest scores in five categories: Marketing Data Reporting via MCP, Historical Performance Insights & Causal Explanation via MCP, Campaign & Ad Set-Level Optimization with MCP, MCP interoperability and workflows and Enterprise-Grade Platform. Sellforte is best for retailers and ecommerce businesses looking for enterprise-grade MMM and incrementality testing, and granular campaign & ad set level optimization.
- Recast: 31.4/48. Second overall, with the highest scores for Channel-Level Optimization with MCP and Incrementality Testing with MCP. It also ties Triple Whale for the highest score in Analytical Backbone.
- Triple Whale: 30.6/48. Third overall and tied with Recast for the highest score in Analytical Backbone. It ranks second for Marketing Data Reporting via MCP and third for Channel-Level Optimization with MCP.
- Lifesight: 29.8/48. Fourth overall, with the second-highest scores for Campaign & Ad Set-Level Optimization with MCP and Enterprise-Grade Platform. Its highest normalized category scores are in Historical Performance Insights & Causal Explanation via MCP, followed by Channel-Level Optimization with MCP.
- Measured: 26.6/48. Fifth overall and second for Incrementality Testing with MCP. Its highest normalized category scores are in Marketing Data Reporting via MCP and Historical Performance Insights & Causal Explanation via MCP, followed by Analytical Backbone.
- Fospha: 24.3/48. Sixth overall. Its highest normalized category scores are in MCP interoperability and workflows, followed by Marketing Data Reporting via MCP and Channel-Level Optimization with MCP.
- Haus: 18.9/48. Seventh overall. Its highest normalized category scores are in Incrementality Testing with MCP and Marketing Data Reporting via MCP, followed by Analytical Backbone.
- Northbeam: 17.6/48. Eighth overall. Its highest normalized category scores are in Marketing Data Reporting via MCP and MCP interoperability and workflows.
Introduction and Table of Contents
In 2026, Marketing Mix Modeling (MMM) and incrementality testing vendors have started adding Model Context Protocol (MCP) servers to their solutions, with launches from Measured and Recast among the examples. These connections let marketers bring their measurement data into AI assistants such as Claude and ChatGPT, so they can ask questions and work with results in the tools they already use.
The capabilities available through these connections differ. A marketer might want to retrieve performance data, understand what drove sales, calculate a new budget scenario or review an incrementality test. The depth of support for each task varies by vendor, as does the level of detail available for channels, campaigns and ad sets. Buyers need to understand what each solution makes available through MCP and how well it supports their day-to-day decisions.
We conducted this evaluation to bring greater transparency to the market. Using a common framework of 48 criteria across eight categories, Claude and ChatGPT assessed public documentation for eight vendors. By publishing the criteria, scoring method and results, we aim to make the differences easier to compare and help marketers identify the capabilities they should examine more closely when choosing a solution.
- What is an MCP server for MMM and incrementality testing?
- Research methodology: evaluation criteria
- Research methodology: scoring
- Which vendors were evaluated?
- In-depth comparison by category
- Summary by vendor
- Frequently asked questions
- Change log
- Evaluation dates by vendor and LLM
- Limitations and disclosures
- Further reading and resources
What Is an MCP Server for MMM and Incrementality Testing?
An MCP server for Marketing Mix Modeling (MMM) and incrementality testing enables marketers to interact with their marketing measurement solution through an AI assistant such as Claude or ChatGPT. They can ask questions in plain language and work with their company’s measurement data and supported analytical tools directly in the conversation.
For MMM, this can mean asking which channels drove incremental sales, how campaign-level incremental ROAS compares with platform-reported ROAS, or how a different budget allocation would affect expected revenue. For incrementality testing, it can mean retrieving the results of a completed geo test, reviewing measured lift or identifying which experiment to run next, depending on the capabilities the vendor makes available.
Through the MCP connection, the assistant requests data or calls a supported tool in the measurement solution. The solution returns the relevant model outputs, test results or scenario calculations, which the assistant can explain and use in further analysis. The functions available through MCP vary by vendor, which is why this evaluation examines individual capabilities.
Research Methodology: Evaluation Criteria
This article follows Sellforte's Solution Research Objectives and Guiding principles: We first conducted primary research to form an evaluation framework. We then gave that framework to Claude and ChatGPT for conducting the vendor evaluation.
To develop the 48 criteria in the evaluation framework on this article, we:
- Anayzed 5,776 prompts that marketers used in interacting with Sellforte via MCP or Sellforte's in-platform conversational AI interface. These prompts gave us an understanding of the real usage patterns and expectations that marketers have towards MCPs.
- Analyzed 570 discussions with marketers, where MCP and conversational AI interfaces in MMM and incrementality testing was discussed.
- Reviewed our experience in more than 20 retail and ecommerce RFPs and tender processes that contained requirements for MCPs and AI capabilities of MMM and incrementality testing platforms.
- Used desk research and LLM-assisted investigation to identify gaps and check the framework against public information.
Category 1. Marketing Data Reporting via MCP
These criteria assess access to actual business and media data: online and physical-store sales, digital and offline media, additional business outcomes, and filtering by the dimensions used in planning. Reporting access is the starting point for a useful marketing assistant.
The criteria cover both the breadth of available data and the ability to retrieve a relevant slice of it. For example, a marketer might request sales and media spend for a particular market and date.
| ID | Criterion | What it means |
|---|---|---|
| 1.1 | MCP reports sales progress for online sales | MCP reports actual online sales for a specified period. |
| 1.2 | MCP reports sales progress for offline store sales | MCP reports actual physical-store sales for a specified period, using offline store sales data retrieved from customer's data warehouse. |
| 1.3 | MCP reports digital media data (spend, media metrics) | MCP reports digital media data across Meta, Google, TikTok, and other major paid platforms. |
| 1.4 | MCP reports offline media data (spend, media metrics) | MCP reports spend and media metric data for offline media, covering at least TV, out-of-home, radio, print. |
| 1.5 | MCP reports business outcomes beyond revenue | MCP reports supported business outcomes such as contribution margin, orders or new customers. |
| 1.6 | Filters and groups by business dimensions | MCP applies explicit date, brand, market, product and campaign filters and returns results at the requested supported granularity. |
Category 2. Historical Performance Insights & Causal Explanation via MCP
This category examines access to modeled incremental returns, promotion effects and the drivers of performance. The criteria distinguish digital and offline channel contributions and check whether MMM-based incremental ROAS is updated daily. Decomposing results into baseline, media and non-media contributions helps marketers understand which modeled drivers explain a change in sales over a given period.
| ID | Criterion | What it means |
|---|---|---|
| 2.1 | MCP reports incremental ROAS and incremental revenue for each digital channel | MCP reports true incremental impact, not just last-click or platform-reported ROAS, for each digital channel. |
| 2.2 | MCP reports incremental ROAS and incremental revenue for each offline channel | MCP reports incremental ROAS and revenue for offline channels such as TV, OOH, and radio. |
| 2.3 | MCP reports promotion-driven revenue, in addition to media-driven | MCP surfaces how promotions and pricing changes contributed to sales, not just paid media. |
| 2.4 | MCP's incremental ROAS measurement is updated daily (not weekly or monthly), based on MMM | MCP provides daily measurement of incremental ROAS based on MMM, not just weekly, monthly or quarterly model refreshes. |
| 2.5 | Decomposes performance drivers (Base, media, non-media drivers) | MCP exposes modeled contributions from media, baseline and supported non-media drivers to explain changes in business outcomes. |
Category 3. Channel-Level Optimization with MCP
These criteria follow a planning workflow from a proposed budget to an optimized allocation, forecast, constraints, scenario comparison and saved plan. The criteria also cover response curves and marginal returns, which inform where additional spend could contribute most. A marketer should be able to explore budget changes against a business target while retaining channel limits and planning dates.
| ID | Criterion | What it means |
|---|---|---|
| 3.1 | MCP creates an optimized channel allocation | Invokes the provider’s optimization engine through MCP to calculate channel budgets for a specified objective and period. |
| 3.2 | MCP forecasts a specified media plan | Invokes the provider’s forecasting engine through MCP to estimate outcomes for a specified budget allocation. |
| 3.3 | MCP reports marginal returns and response curves | Retrieves marginal returns and spend-response data for supported channels through MCP. |
| 3.4 | MCP simulates a user-specified budget change | Calculates the modeled impact of a specified change to channel spend through MCP. |
| 3.5 | MCP enforces planning constraints | Passes user-defined budget limits, channel restrictions and planning dates to the optimization engine and exposes the resulting constraints. |
| 3.6 | MCP compares scenarios with a baseline plan | Returns comparable spend and outcome estimates for alternative scenarios and an explicit baseline plan through MCP. |
| 3.7 | MCP plans against a business target | Calculates a budget plan for a supported target such as incremental profit, customer acquisition or a specified outcome level. |
| 3.8 | MCP retrieves saved plans and assumptions | Retrieves previously saved scenarios through MCP with their identifiers, inputs and outputs. |
Category 4. Campaign & Ad Set-Level Optimization with MCP
Campaign and ad-set decisions need more granular evidence than channel allocation. The criteria separately assess incremental and marginal returns, budget and bid recommendations, execution through ad-platform APIs, and post-change reporting. The comparison with last-click and ad-platform ROAS helps put modeled incremental returns in context. Separate criteria distinguish receiving a recommendation from applying it, then reviewing revenue and spend before and after the change.
| ID | Criterion | What it means |
|---|---|---|
| 4.1 | MCP reports incremental ROAS for each campaign and ad set | MCP reports incremental revenue and ROAS at the individual campaign and ad set level, not just at the channel level. |
| 4.2 | For each campaign and ad set, MCP compares incremental ROAS to last-click and ad platform attribution ROAS | Returns comparable incremental, last-click and ad-platform returns for supported campaigns and ad sets. |
| 4.3 | For each campaign and ad set, MCP provides marginal returns (miROAS) | Retrieves modeled marginal returns for individual campaigns and ad sets through MCP. |
| 4.4 | For each campaign and ad set, MCP recommends optimal daily spend/budget | Returns daily budget recommendations through MCP. |
| 4.5 | For each campaign and ad set, MCP recommends optimal bid value (e.g., Target ROAS) | Returns bidding targets (e.g., Target ROAS) through MCP. |
| 4.6 | MCP can push daily spend/budget changes to Meta, Google etc. APIs | MCP can execute daily spend/budget changes directly on major ad platforms via API. |
| 4.7 | MCP can push bidding changes (e.g., Target ROAS) to Meta, Google etc. APIs | MCP can execute bidding parameter changes (e.g. Target ROAS) directly on major ad platforms via API. |
| 4.8 | For each executed bidding change for a campaign or ad set, MCP provides pre/post analysis summarizing revenue and spend impact of the change | After a bidding or budget change is applied, MCP can provide the actual impact using a pre/post comparison for bidding change at the campaign & ad set level. |
Category 5. Incrementality Testing with MCP
The testing category evaluates whether the there's MCP-based access to geo tests, owned-media A/B tests, and conversion platform-lift results. It also evaluates whether the MCP is capable of prioritizing and designing new experiments.
| ID | Criterion | What it means |
|---|---|---|
| 5.1 | MCP reports results for Geo Tests | Retrieves completed geographic incrementality-test results with the tested intervention and outcome identified. |
| 5.2 | MCP reports results for Own Media A/B tests (e.g. leaflet tests) | Retrieves incrementality results for randomized or controlled owned-media tests such as email or leaflet experiments. |
| 5.3 | MCP reports findings for Meta Conversion Lift tests | Retrieves incrementality results from advertising-platform lift studies such as Meta Conversion Lift. |
| 5.4 | MCP prioritizes experiments | Uses available measurement results and uncertainty to recommend specific hypotheses, channels or markets for testing. |
| 5.5 | MCP designs feasible incrementality tests | Produces a test design specifying treatment and control, power or minimum detectable effect, duration and required spend. |
Category 6. MCP interoperability and workflows
This category assess MCP's connectivity, user experience and ability to embed it in agentic workflows. The category considers how marketers connect their assistant, interpret the available tools and metrics, and reuse results in spreadsheets or reports. Criteria include setup guidance, server-provided visual outputs, tool and metric documentation, reusable structured data, and recurring workflows in external agents.
| ID | Criterion | What it means |
|---|---|---|
| 6.1 | Documented instructions for connecting with Claude and ChatGPT | Provides documented MCP connection and authentication instructions for Claude and ChatGPT |
| 6.2 | MCP provides native visual outputs | Returns server-provided tables or chart artifacts that a documented compatible MCP client can display. |
| 6.3 | Exposes tool and metric documentation | Makes available discoverable MCP tool definitions and documentation for interpreting the returned data. |
| 6.4 | Returns reusable structured data | Provides complete structured results or downloadable data files for use in spreadsheets, reports and downstream tools. |
| 6.5 | Supports external agent automation | Documents repeatable MCP invocation from an external agent runtime for recurring analysis workflows. |
Category 7. Analytical Backbone
This category assesses the measurement system behind the answer: a Bayesian MMM foundation, experiment calibration, model validation, editable settings and controlled calibration through MCP. The criteria examine whether recommendations come from the underlying model and whether users can inspect evidence of its quality, such as validation metrics. They also cover how experiments inform calibration and whether configuration changes can be inspected, made with appropriate authorization and tracked through version history.
| ID | Criterion | What it means |
|---|---|---|
| 7.1 | MCP provides deterministic, model-backed answers with Bayesian MMM as backbone | MCP-provided recommendations are grounded in a Bayesian Marketing Mix Model, not last-click attribution or descriptive analytics. |
| 7.2 | The underlying MMM used by the MCP is calibrated with incrementality tests | The MMM is calibrated against real incrementality test results, priors and posteriors informed by causal experiments. |
| 7.3 | MCP reports model validation and other modelling KPIs | The tool reports model validation metrics (R2, MAPE, posterior predictive checks, holdout performance) so users can assess model quality. |
| 7.4 | Model calibration & configuration settings (e.g., priors) are auditable and editable in a self-serve UI | Customers can inspect and configure model priors and other key parameters in a self-serve UI, not just accept the model as a black box. |
| 7.5 | Supports controlled calibration through MCP | Allows authorized agents to propose or apply model-calibration changes through MCP with validation and version history. |
Category 8. Enterprise-Grade Platform
These criteria cover the platforms' ability to cater to the customer segment with highest demands for the platform: enterprises. Criteria include enterprise references, independent security assurance, residency and cloud options, single sign-on, and public hands-on MCP access.
| ID | Criterion | What it means |
|---|---|---|
| 8.1 | At least 10 public reference customers from $1B+ revenue brands | Proven track record with large, sophisticated advertisers, not just mid-market or DTC brands. |
| 8.2 | SOC 2, ISO 27001, or audited IT security by a third-party cyber security auditor | Independently verified security posture, a baseline requirement for enterprise IT procurement. |
| 8.3 | Data residency: geography option between US and EU | Customers can choose between US and EU data-residency options for the relevant service. |
| 8.4 | Multi-cloud: option between AWS, GCP, and Azure | Deployment flexibility to match the customer's existing cloud infrastructure. |
| 8.5 | Supports single sign-on (SSO) for enterprises | Supports enterprise single sign-on for customer authentication. |
| 8.6 | Hands-on demo or trial of the MCP capabilities is available without sales-call gating | Provides a hands-on, promptable MCP demo or trial without requiring a sales call. |
Research Methodology: Scoring
Scoring each tool against the evaluation criteria was done by Claude and ChatGPT. See specific LLM-model versions in Evaluation Dates by Vendor and LLM section. Here were the assessment steps:
Step 1. Both LLMs independently scored each criterion for each vendor, based on the instructions in this section.
Step 2. For the final score for each criteria, we used an average between the two LLMs, rounded to 1 decimal.
Step 3. Category scores were created by summing up the scores from each criteria in the category, and total score was calculated by by summing up the category scores.
The full research notes include the evaluation criteria, instructions and individual Claude and ChatGPT evaluations for each vendor.
Why is the scoring done by Claude and ChatGPT?
Simulating evaluation a buyer might make based on public materials. By using only vendor-provided materials on their website and technical documentation in the assessment, Claude and ChatGPT -based evaluation simulates how a buyer without prior knowledge of the vendor might assess the vendor prior to a sales call or demo meeting.
Transparent methodology and reproducible assessment. Because the evaluation instructions are published alongside this article, any reader can re-run the evaluation for any vendor and verify or challenge the results. This is a higher standard of transparency than conventional analyst-style research, where the scoring rationale is typically not disclosed.
Equal treatment of vendors. Claude and ChatGPT apply the same instructions, the same criteria, the same scoring scale, and the same source prioritization rules to every vendor in the comparison.
Instructions given to Claude and ChatGPT
The following instructions were given to Claude and ChatGPT for conducting the evaluation.
| ID | Instruction |
|---|---|
| 1 | You are an independent evaluator of MCP servers for Marketing Mix Modeling and Incrementality testing |
| 2 | Use these evaluation criteria for categories 1-5: 1.00: Strong public evidence that the capability exists in the MCP 0.75: Partial public evidence that the capability exists in the MCP 0.50: No evidence found that the capability is available via the MCP, but there is strong evidence that the capability is available in the broader platform 0.25: No evidence found that the capability is available via the MCP, but there is partial evidence that the capability is available in the broader platform 0.00 : No evidence found that the capability exists in the MCP or in the broader platform Use these evaluation criteria for categories 6-8: 1: Strong evidence that the platform supports the capability 0.5: Partial evidence that the platform supports the capability 0: No evidence found that the platform supports the capability |
| 3 | In the evaluation, only use information available at the company website, and in the domain where technical documentation is located (if separately hosted). |
| 4 | Prioritize type of source materials in this order 1. Technical documentation (such as support center) 2. Product page 3. Marketing collateral (such as product launch blog posts) |
| 5 | When interpreting terminology, you can assume that - ROI is the same thing as incremental ROAS or iROAS - Marginal ROI is the same thing as Marginal Incremental ROAS or miROAS - Diminishing return curves are the same thing as response curves |
| 6 | For the purposes of this evaluation, feature available via API also qualify with same scores as features via MCP. |
| 7 | As an output, provide your evaluation in an excel file, using these columns: - Category ID - Category - Criteria ID - Criteria - Criteria Score - Two-sentence rationale for the score. If you found no evidence, comment that you did not find evidence, instead of claiming that the capability does not exist - URL to source - Type of source (Technical doc, Product page, Marketing collateral) |
As you can see from the scoring model, the LLMs are instructed to evaluate publicly available documentation. This means that low scores or zero scores may reflect the absence of publicly documentation, not necessarily absence of the capability itself.
Why do different categories use different scoring scales?
Categories 1–5 assess the measurement and optimization capabilities that marketers can access through the MCP. Scores 1.0 or 0.75 are used if there is strong or partial evidence that the capability is available via the MCP.
We also asked the LLMs to give a score of 0.5 or 0.25 if there is evidence for a capability being available in the overall platform, but there is no evidence for it being available via the MCP. As an example, a platform might be able to show geo test results, but there is no documentation where the MCP can summarize their findings. Why did we apply make this decision? Firstly, it is possible that the capability is already available via the MCP, but the documentation does not yet exist. Secondly, if the capability is not in MCP yet, it is likely that it will be added to the MCP in the near future.
Categories 6–8 assess MCP interoperability and workflows, Analytical Backbone, and Enterprise-Grade Platform. For these criteria, the question is whether there is evidence for the overall solution supporting the specified requirements. Thus, a simpler three-level scale was used: 0, 0.5. 1.0.
Why do API capabilities qualify for the same scores as MCP capabilities?
We included equivalent API capabilities because the evaluation aims to assess the measurement data and functions that can support AI-assisted marketing workflows. An MCP tool can call an existing API to retrieve data or run an operation, as described in the official MCP tools specification. A documented API can therefore provide a route to the same underlying capability through an integration. Giving it equivalent credit allows us to recognize that access even when the vendor has not exposed the capability through its own MCP server.
This is a scoring choice for this research. An API must provide evidence of the capability being assessed; simply offering an API is not enough. API-supported functionality may require additional integration work before a marketer can use it through an AI assistant, so an equivalent score does not imply an identical setup experience.
How to reproduce the evaluation and scoring yourself with Claude or ChatGPT?
To reproduce the evaluation for any vendor in this study, you can follow the instructions in this section.
Step 1. Open Claude or ChatGPT incognito mode. Choose the model and level of effort.
Step 2. Provide it following context:
-
Evaluation instructions that have been shared in this article
-
Evaluation criteria that have been shared in this article
Step 3. Initiate the prompt: "Evaluate [add vendor name, e.g., Sellforte], based on the instructions and criteria provided as context for this prompt."
What biases and limitations does our scoring approach have, and how are we addressing them?
While LLM-based evaluation reduces vendor bias and promotes equal treatment of vendors and reproducibility, there are four main limitations to be aware of.
1. Availability and of documentation that each vendor has made public. A vendor with extensive and detailed public documentation will naturally score higher than one that keeps product details behind a sales call gate, even if the underlying capabilities are comparable. Each vendor can affect their own scoring through public documentation.
2. LLMs' ability to find public documentation. Even with web search enabled, Claude and ChatGPT may not find every relevant page on a vendor's website. Documentation that is not well-indexed or that sits in obscure subdomains may be missed. We tried to address this by using advanced models from two separate LLMs, and added a specific point to the instructions to search for vendor-provided technical documentation that might sometimes be under a different sub-domain.
3. LLMs' ability to interpret public documentation against the scoring criteria. Matching a product description to a specific criterion requires judgment. We reduce interpretation variance by using advanced models from two separate LLMs, providing detailed criterion descriptions, and providing explicit terminology equivalences. However, edge cases may still exist.
4. Limitations of LLM technology, including hallucinations. LLMs are known to occasionally assert things that are not based on facts. To reduce this risk, we used advanced models from two separate LLMs, asked them to provide source URL for their assessment, and asked them to provide a rationale for the score in each criterion.
Which MCP Vendors Were Evaluated?
We focused this article on MCP servers from recgonized vendors that provide Marketing Mix Modeling and/or incrementality testing. When searching for vendors that would meet these criteria we looked into multiple product categories: Data connector companies, traditional Marketing Mix Modeling providers, next gen MMM vendors, Incrementality testing tools, attribution tools.
We ended up evaluating eight MCP servers for MMM and Incrementality Testing: Sellforte, Recast, Triple Whale, Lifesight, Measured Fospha, Haus, Northbeam.
In-depth Comparison of MCP Servers for MMM and Incrementality Testing
The charts and tables in this section show the research scores and the results for all research categories
For the individual Claude and ChatGPT evaluations behind these results, see the full research notes.
Total research score across all 48 criteria
Claude and ChatGPT evaluations dated September 23, 2026. Bars and ranking use unrounded scores. Labels are rounded to one decimal.
Sellforte scored 38.3/48, followed by Recast at 31.4, Triple Whale at 30.6, Lifesight at 29.8, Measured at 26.6, Fospha at 24.3, Haus at 18.9, and Northbeam at 17.6.
The average across the eight vendors is 27.2/48. The three categories with the greatest variation in vendor scores were
-
Channel-Level Optimization with MCP (score range 2.3 - 7.9 out of 8)
-
Historical Performance Insights & Causal Explanation via MCP (score range 1.6 - 5.0 out of 5)
-
MCP interoperability and workflows (score ranged 1.3 - 4.3 out of 5)
| Evaluated category | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Max |
|---|---|---|---|---|---|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 6.0 | 5.0 | 5.3 | 4.1 | 4.9 | 4.3 | 3.1 | 5.0 | 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 5.0 | 4.1 | 3.6 | 3.8 | 3.6 | 2.4 | 1.9 | 1.6 | 5 |
| 3. Channel-Level Optimization with MCP | 7.4 | 7.9 | 6.5 | 5.6 | 4.5 | 4.6 | 2.9 | 2.3 | 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 5.0 | 1.0 | 3.8 | 4.1 | 2.4 | 3.3 | 3.1 | 1.4 | 8 |
| 5. Incrementality Testing with MCP | 2.6 | 3.6 | 2.5 | 2.1 | 3.0 | 1.3 | 2.6 | 1.1 | 5 |
| 6. MCP interoperability and workflows | 4.3 | 4.0 | 3.8 | 3.5 | 3.0 | 3.8 | 1.3 | 3.5 | 5 |
| 7. Analytical Backbone | 3.5 | 3.8 | 3.8 | 3.5 | 3.3 | 2.0 | 2.3 | 1.5 | 5 |
| 8. Enterprise-Grade Platform | 4.5 | 2.0 | 1.5 | 3.0 | 2.0 | 2.8 | 1.8 | 1.3 | 6 |
| Total | 38.3 | 31.4 | 30.6 | 29.8 | 26.6 | 24.3 | 18.9 | 17.6 | 48 |
1. Marketing Data Reporting via MCP
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The six criteria assess reporting through MCP: online and offline store sales, digital and offline media data, business outcomes beyond revenue, and filtering or grouping results by business dimensions.
What the research found. The research scores were: Sellforte (6.0/6), Triple Whale (5.3/6), Recast (5.0/6), Northbeam (5.0/6), Measured (4.9/6), Fospha (4.3/6), Lifesight (4.1/6), Haus (3.1/6). Recast and Northbeam are tied.
Highest average research scores were for reporting business outcomes beyond revenue and digital media data. Lower average research scores were for offline store sales and offline media reporting. These two reporting criteria also showed the greatest variation among vendors, with offline store sales varying most.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 1.1 MCP reports sales progress for online sales | 1.0 | 0.9 | 1.0 | 0.6 | 0.8 | 1.0 | 0.6 | 1.0 | 0.9 |
| 1.2 MCP reports sales progress for offline store sales | 1.0 | 0.9 | 0.8 | 0.5 | 0.8 | 0.1 | 0.4 | 0.5 | 0.6 |
| 1.3 MCP reports digital media data (spend, media metrics) | 1.0 | 0.8 | 1.0 | 0.9 | 1.0 | 1.0 | 0.8 | 1.0 | 0.9 |
| 1.4 MCP reports offline media data (spend, media metrics) | 1.0 | 0.8 | 0.6 | 0.5 | 0.8 | 0.3 | 0.4 | 0.8 | 0.6 |
| 1.5 MCP Reports business outcomes beyond revenue | 1.0 | 1.0 | 1.0 | 0.9 | 0.9 | 1.0 | 0.8 | 1.0 | 0.9 |
| 1.6 Filters and groups by business dimensions | 1.0 | 0.8 | 0.9 | 0.8 | 0.8 | 0.9 | 0.3 | 0.8 | 0.8 |
| Category total | 6.0 | 5.0 | 5.3 | 4.1 | 4.9 | 4.3 | 3.1 | 5.0 | 4.7 |
2. Historical Performance Insights & Causal Explanation via MCP
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The five criteria assess historical performance insights through MCP, including incremental ROAS and revenue for digital and offline channels, promotion-driven revenue, daily MMM-based measurement updates, and the decomposition of base, media and non-media performance drivers.
What the research found. The research scores were: Sellforte (5.0/5), Recast (4.1/5), Lifesight (3.8/5), Triple Whale (3.6/5), Measured (3.6/5), Fospha (2.4/5), Haus (1.9/5), Northbeam (1.6/5). Triple Whale and Measured are tied.
Highest average research scores were for reporting incremental ROAS and revenue for digital channels and decomposing performance drivers. Lower average research scores were for daily MMM-based updates to incremental ROAS. Promotion-driven revenue reporting and daily MMM-based updates showed the greatest variation among vendors.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 2.1 MCP reports incremental ROAS and incremental revenue for each digital channel | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 0.8 | 0.5 | 0.9 |
| 2.2 MCP reports incremental ROAS and incremental revenue for each offline channel | 1.0 | 1.0 | 0.6 | 0.6 | 0.8 | 0.1 | 0.5 | 0.3 | 0.6 |
| 2.3 MCP reports promotion-driven revenue, in addition to media-driven | 1.0 | 1.0 | 1.0 | 0.8 | 0.5 | 0.1 | 0.3 | 0.4 | 0.6 |
| 2.4 MCP's incremental ROAS measurement is updated daily (not weekly or monthly), based on MMM | 1.0 | 0.1 | 0.1 | 0.5 | 0.6 | 0.9 | 0.1 | 0.3 | 0.5 |
| 2.5 Decomposes performance drivers (Base, media, non-media drivers) | 1.0 | 1.0 | 0.9 | 0.9 | 0.8 | 0.3 | 0.3 | 0.3 | 0.7 |
| Category total | 5.0 | 4.1 | 3.6 | 3.8 | 3.6 | 2.4 | 1.9 | 1.6 | 3.3 |
3. Channel-Level Optimization with MCP
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The eight criteria assess channel-level planning through MCP: creating optimized allocations, forecasting media plans, reporting marginal returns and response curves, simulating budget changes, applying planning constraints, comparing scenarios, planning against business targets, and retrieving saved plans and assumptions.
What the research found. The research scores were: Recast (7.9/8), Sellforte (7.4/8), Triple Whale (6.5/8), Lifesight (5.6/8), Fospha (4.6/8), Measured (4.5/8), Haus (2.9/8), Northbeam (2.3/8).
Highest average research scores were for simulating a user-specified budget change and forecasting a media plan. Lower average research scores were for enforcing planning constraints and retrieving saved plans and assumptions, which tied for the lowest average. Those two criteria also showed the greatest variation among vendors, with retrieval of saved plans and assumptions varying most.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 3.1 MCP creates an optimized channel allocation | 1.0 | 1.0 | 0.8 | 0.8 | 0.5 | 0.5 | 0.5 | 0.5 | 0.7 |
| 3.2 MCP forecasts a specified media plan | 1.0 | 1.0 | 0.8 | 0.8 | 0.5 | 0.8 | 0.5 | 0.5 | 0.7 |
| 3.3 MCP reports marginal returns and response curves | 0.5 | 1.0 | 0.8 | 0.8 | 0.8 | 0.9 | 0.6 | 0.3 | 0.7 |
| 3.4 MCP simulates a user-specified budget change | 1.0 | 1.0 | 0.8 | 0.8 | 0.5 | 1.0 | 0.5 | 0.5 | 0.8 |
| 3.5 MCP enforces planning constraints | 1.0 | 1.0 | 0.8 | 0.5 | 0.5 | 0.5 | 0.3 | 0.0 | 0.6 |
| 3.6 MCP compares scenarios with a baseline plan | 1.0 | 0.9 | 1.0 | 0.8 | 0.6 | 0.4 | 0.3 | 0.3 | 0.6 |
| 3.7 MCP plans against a business target | 0.9 | 1.0 | 0.8 | 0.8 | 0.5 | 0.4 | 0.3 | 0.3 | 0.6 |
| 3.8 MCP retrieves saved plans and assumptions | 1.0 | 1.0 | 1.0 | 0.6 | 0.6 | 0.3 | 0.0 | 0.0 | 0.6 |
| Category total | 7.4 | 7.9 | 6.5 | 5.6 | 4.5 | 4.6 | 2.9 | 2.3 | 5.2 |
4. Campaign & Ad Set-Level Optimization with MCP
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The eight criteria assess campaign and ad set optimization through MCP, from reporting incremental ROAS and comparing it with attribution-based ROAS to providing marginal returns, recommending daily budgets and bid values, pushing budget and bidding changes to ad platforms, and analyzing the revenue and spend impact of executed bidding changes.
What the research found. The research scores were: Sellforte (5.0/8), Lifesight (4.1/8), Triple Whale (3.8/8), Fospha (3.3/8), Haus (3.1/8), Measured (2.4/8), Northbeam (1.4/8), Recast (1.0/8).
Highest average research scores were for comparing incremental ROAS with last-click and ad platform attribution ROAS, and reporting incremental ROAS by campaign and ad set. Lower average research scores were for recommending optimal bid values and pushing bidding changes to ad platforms, which tied for the lowest average. Campaign and ad set incremental ROAS reporting and its comparison with attribution-based ROAS showed the greatest variation among vendors.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 4.1 MCP reports incremental ROAS for each campaign and ad set | 1.0 | 0.3 | 0.8 | 0.9 | 0.8 | 0.8 | 0.8 | 0.3 | 0.7 |
| 4.2 For each campaign and ad set, MCP compares incremental ROAS to last-click and ad platform attribution ROAS | 1.0 | 0.1 | 0.8 | 0.8 | 0.8 | 0.9 | 0.5 | 0.8 | 0.7 |
| 4.3 For each campaign and ad set, MCP provides marginal returns (miROAS) | 0.5 | 0.0 | 0.5 | 0.5 | 0.4 | 0.1 | 0.3 | 0.0 | 0.3 |
| 4.4 For each campaign and ad set, MCP recommends optimal daily spend/budget | 0.5 | 0.5 | 0.4 | 0.5 | 0.3 | 0.4 | 0.4 | 0.1 | 0.4 |
| 4.5 For each campaign and ad set, MCP recommends optimal bid value (e.g., Target ROAS) | 0.5 | 0.0 | 0.1 | 0.3 | 0.0 | 0.1 | 0.3 | 0.3 | 0.2 |
| 4.6 MCP can push daily spend/budget changes to Meta, Google etc. APIs | 0.5 | 0.0 | 0.5 | 0.5 | 0.1 | 0.4 | 0.6 | 0.0 | 0.3 |
| 4.7 MCP can push bidding changes (e.g., Target ROAS) to Meta, Google etc. APIs | 0.5 | 0.0 | 0.3 | 0.5 | 0.0 | 0.1 | 0.1 | 0.0 | 0.2 |
| 4.8 For each executed bidding change for a campaign or ad set, MCP provides pre/post analysis summarizing revenue and spend impact of the change | 0.5 | 0.1 | 0.5 | 0.3 | 0.1 | 0.5 | 0.3 | 0.0 | 0.3 |
| Category total | 5.0 | 1.0 | 3.8 | 4.1 | 2.4 | 3.3 | 3.1 | 1.4 | 3.0 |
5. Incrementality Testing with MCP
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The five criteria assess incrementality testing through MCP: reporting results from geo tests, own-media A/B tests and Meta Conversion Lift tests, prioritizing experiments, and designing feasible incrementality tests.
What the research found. The research scores were: Recast (3.6/5), Measured (3.0/5), Sellforte (2.6/5), Haus (2.6/5), Triple Whale (2.5/5), Lifesight (2.1/5), Fospha (1.3/5), Northbeam (1.1/5). Sellforte and Haus are tied.
Highest average research scores were for reporting geo-test results and prioritizing experiments. Lower average research scores were for reporting own-media A/B test results and Meta Conversion Lift findings, which tied for the lowest average. Those two reporting criteria also showed the greatest variation among vendors, with own-media A/B test reporting varying most.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 5.1 MCP reports results for Geo Tests | 0.8 | 0.9 | 1.0 | 0.8 | 0.9 | 0.5 | 0.8 | 0.5 | 0.8 |
| 5.2 MCP reports results for Own Media A/B tests (e.g. leaflet tests) | 0.8 | 0.9 | 0.0 | 0.3 | 0.5 | 0.0 | 0.6 | 0.0 | 0.4 |
| 5.3 MCP reports findings for Meta Conversion Lift tests | 0.8 | 0.9 | 0.6 | 0.1 | 0.1 | 0.3 | 0.3 | 0.0 | 0.4 |
| 5.4 MCP prioritizes experiments | 0.1 | 0.5 | 0.4 | 0.5 | 1.0 | 0.3 | 0.5 | 0.3 | 0.4 |
| 5.5 MCP designs feasible incrementality tests | 0.3 | 0.5 | 0.5 | 0.5 | 0.5 | 0.3 | 0.5 | 0.4 | 0.4 |
| Category total | 2.6 | 3.6 | 2.5 | 2.1 | 3.0 | 1.3 | 2.6 | 1.1 | 2.4 |
6. MCP interoperability and workflows
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The five criteria assess MCP interoperability and workflows: documented instructions for connecting with Claude and ChatGPT, native visual outputs, accessible tool and metric documentation, reusable structured data, and automation through external agents.
What the research found. The research scores were: Sellforte (4.3/5), Recast (4.0/5), Triple Whale (3.8/5), Fospha (3.8/5), Lifesight (3.5/5), Northbeam (3.5/5), Measured (3.0/5), Haus (1.3/5). Triple Whale and Fospha are tied, as are Lifesight and Northbeam.
Highest average research scores were for returning reusable structured data and providing connection instructions for Claude and ChatGPT. Lower average research scores were for native visual outputs. Native visual outputs and external agent automation showed the greatest variation among vendors.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 6.1 Documented instructions for connecting with Claude and ChatGPT | 0.8 | 1.0 | 1.0 | 1.0 | 0.8 | 1.0 | 0.3 | 1.0 | 0.8 |
| 6.2 MCP provides native visual outputs | 1.0 | 0.0 | 0.0 | 0.5 | 0.5 | 0.0 | 0.0 | 0.0 | 0.3 |
| 6.3 Exposes tool and metric documentation | 1.0 | 1.0 | 0.8 | 0.8 | 0.3 | 1.0 | 0.5 | 0.8 | 0.8 |
| 6.4 Returns reusable structured data | 1.0 | 1.0 | 1.0 | 0.8 | 1.0 | 1.0 | 0.5 | 1.0 | 0.9 |
| 6.5 Supports external agent automation | 0.5 | 1.0 | 1.0 | 0.5 | 0.5 | 0.8 | 0.0 | 0.8 | 0.6 |
| Category total | 4.3 | 4.0 | 3.8 | 3.5 | 3.0 | 3.8 | 1.3 | 3.5 | 3.4 |
7. Analytical Backbone
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The five criteria assess the analytical foundations behind MCP responses: deterministic answers backed by Bayesian MMM, calibration with incrementality tests, reporting model validation and modelling KPIs, auditable and editable calibration settings in a self-serve interface, and controlled model calibration through MCP.
What the research found. The research scores were: Recast (3.8/5), Triple Whale (3.8/5), Sellforte (3.5/5), Lifesight (3.5/5), Measured (3.3/5), Haus (2.3/5), Fospha (2.0/5), Northbeam (1.5/5). Recast and Triple Whale are tied, as are Sellforte and Lifesight.
Highest average research scores were for calibrating the underlying MMM with incrementality tests and reporting model validation and modelling KPIs. Lower average research scores were for controlled calibration through MCP. Model validation reporting and auditable, editable calibration settings in a self-serve interface showed the greatest variation among vendors.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 7.1 MCP provides deterministic, model-backed answers with Bayesian MMM as backbone | 1.0 | 1.0 | 0.8 | 0.5 | 0.5 | 0.5 | 0.5 | 0.3 | 0.6 |
| 7.2 The underlying MMM used by the MCP is calibrated with incrementality tests | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 0.5 | 1.0 | 0.8 | 0.9 |
| 7.3 MCP reports model validation and other modelling KPIs | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 0.8 | 0.5 | 0.0 | 0.8 |
| 7.4 Model calibration & configuration settings (e.g., priors) are auditable and editable in a self-serve UI | 0.5 | 0.5 | 1.0 | 1.0 | 0.8 | 0.3 | 0.3 | 0.5 | 0.6 |
| 7.5 Supports controlled calibration through MCP | 0.0 | 0.3 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Category total | 3.5 | 3.8 | 3.8 | 3.5 | 3.3 | 2.0 | 2.3 | 1.5 | 2.9 |
8. Enterprise-Grade Platform
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist.
What we evaluated. The six criteria assess enterprise readiness: at least ten public reference customers with $1 billion or more in revenue, independent security assurance, US and EU data residency options, a choice of AWS, GCP and Azure, enterprise single sign-on, and a hands-on MCP demo or trial available without a sales call.
What the research found. The research scores were: Sellforte (4.5/6), Lifesight (3.0/6), Fospha (2.8/6), Recast (2.0/6), Measured (2.0/6), Haus (1.8/6), Triple Whale (1.5/6), Northbeam (1.3/6). Recast and Measured are tied.
Highest average research scores were for independent security assurance and enterprise single sign-on. Lower average research scores were for access to a hands-on MCP demo or trial without a sales call and a choice of cloud providers. US and EU data residency options and the choice of cloud providers showed the greatest variation among vendors.
| Criterion | Sellforte | Recast | Triple Whale | Lifesight | Measured | Fospha | Haus | Northbeam | Average research score |
|---|---|---|---|---|---|---|---|---|---|
| 8.1 At least 10 public reference customers from $1B+ revenue brands | 0.8 | 0.5 | 0.0 | 0.3 | 0.5 | 0.5 | 0.8 | 0.0 | 0.4 |
| 8.2 SOC 2, ISO 27001, or audited IT security by a third-party cyber security auditor | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 |
| 8.3 Data residency: geography option between US and EU | 0.8 | 0.0 | 0.0 | 1.0 | 0.0 | 0.5 | 0.0 | 0.0 | 0.3 |
| 8.4 Multi-cloud: option between AWS, GCP, and Azure | 1.0 | 0.0 | 0.0 | 0.3 | 0.0 | 0.0 | 0.0 | 0.0 | 0.2 |
| 8.5 Supports single sign-on (SSO) for enterprises | 1.0 | 0.5 | 0.5 | 0.5 | 0.5 | 0.8 | 0.0 | 0.3 | 0.5 |
| 8.6 Hands-on demo or trial of the MCP capabilities is available without sales-call gating | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Category total | 4.5 | 2.0 | 1.5 | 3.0 | 2.0 | 2.8 | 1.8 | 1.3 | 2.3 |
What the Research Suggests About the Market
Average of the recorded Claude and ChatGPT scores from evaluations completed September 23, 2026. Bars and rankings use unrounded scores; labels are rounded to one decimal.
The three strongest areas on average were Marketing Data Reporting via MCP (78.4% of available points), MCP interoperability and workflows (67.5%), and Channel-Level Optimization with MCP (65.0%). These are table stakes capabilities that marketers should expect from a vendor.
The areas with lowest research scores on average were Campaign & Ad Set-Level Optimization with MCP (37.5% of available points), Enterprise-Grade Platform (39.1%), and Incrementality Testing with MCP (47.2%). These are the capabilities where vendors have the most opportunity to improve their offering and differentiate themselves from other vendors.
Summary by Vendor
The profiles below summarize the research score by vendor.
For the individual Claude and ChatGPT evaluations behind these results, see the full research notes.
1. Sellforte (research score: 38.3 out of 48)
Research overview
Sellforte received an average research score of 38.3 out of 48, the highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Sellforte is best for retailers and ecommerce businesses looking for enterprise-grade MMM and incrementality testing, and granular campaign & ad set level optimization.
Research score by category
| Category | Sellforte | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 6.0 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 5.0 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 7.4 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 5.0 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 2.6 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 4.3 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 3.5 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 4.5 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 38.3 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Marketing Data Reporting via MCP: 6.0 / 6. This was above the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting online sales progress, reporting offline store sales progress, and reporting digital media data.
- Historical Performance Insights & Causal Explanation via MCP: 5.0 / 5. This was above the eight-vendor average of 3.3 / 5. Within this category, its highest-scoring criteria included reporting incremental ROAS and revenue for digital channels, reporting incremental ROAS and revenue for offline channels, and reporting promotion-driven revenue.
- Channel-Level Optimization with MCP: 7.4 / 8. This was above the eight-vendor average of 5.2 / 8. Within this category, its highest-scoring criteria included creating an optimized channel allocation, forecasting a media plan, and simulating a user-specified budget change.
Where scores were lower
- Incrementality Testing with MCP: 2.6 / 5. Despite being one of Sellforte’s relatively lower-scoring categories, its score was still above the eight-vendor average of 2.4 / 5.
- Campaign & Ad Set-Level Optimization with MCP: 5.0 / 8. Despite being one of Sellforte’s relatively lower-scoring categories, its score was still above the eight-vendor average of 3.0 / 8.
- Analytical Backbone: 3.5 / 5. Despite being one of Sellforte’s relatively lower-scoring categories, its score was still above the eight-vendor average of 2.9 / 5.
Notable Sellforte Reference Customers
Sellforte lists following companies as examples of public reference customers:
- Fashion Ecommerce: bonprix, Azzas 2154, Represent, Odlo
- Home & Furniture Ecommerce: Finnish Design Shop
- Specialty Ecommerce: FCP Euro, Smartphoto
- Grocery Retail: Lidl
- Fashion Retail: C&A, KIK
- Cosmetics Retail: Douglas
- Sport Retail: Interpsort
- Pet Retail: Fressnapf, Musti Group
- Specialty Retail: Tchibo
- Electronics Retail: Verkkokauppa.com
- Other segments: Telenor (Telecommunications), Paysafe (Payments), eBilet (part of Allegro Group, Events)
Sellforte pricing
Sellforte MCP is included in all Sellforte plans by default. Sellforte pricing starts at $1,900 per month for the Incrementality Testing plan, $2,500 per month for the Incremental attribution plan, $4,500 per month for Full funnel MMM plan, and $8,500 for the Enterprise plan.
2. Recast (research score: 31.4 out of 48)

Research overview
Recast received an average research score of 31.4 out of 48, the second-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Recast | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 5.0 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 4.1 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 7.9 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 1.0 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 3.6 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 4.0 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 3.8 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 2.0 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 31.4 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Channel-Level Optimization with MCP: 7.9 / 8. This was above the eight-vendor average of 5.2 / 8. Within this category, its highest-scoring criteria included creating an optimized channel allocation, forecasting a media plan, and reporting marginal returns and response curves.
- Marketing Data Reporting via MCP: 5.0 / 6. This was above the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting business outcomes beyond revenue, reporting online sales progress, and reporting offline store sales progress.
- Historical Performance Insights & Causal Explanation via MCP: 4.1 / 5. This was above the eight-vendor average of 3.3 / 5. Within this category, its highest-scoring criteria included reporting incremental ROAS and revenue for digital channels, reporting incremental ROAS and revenue for offline channels, and reporting promotion-driven revenue.
Where scores were lower
- Campaign & Ad Set-Level Optimization with MCP: 1.0 / 8. This was below the eight-vendor average of 3.0 / 8.
- Enterprise-Grade Platform: 2.0 / 6. This was below the eight-vendor average of 2.3 / 6.
- Incrementality Testing with MCP: 3.6 / 5. Despite being one of Recast’s relatively lower-scoring categories, its score was still above the eight-vendor average of 2.4 / 5.
3. Triple Whale (research score: 30.6 out of 48)

Research overview
Triple Whale received an average research score of 30.6 out of 48, the third-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Triple Whale | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 5.3 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 3.6 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 6.5 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 3.8 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 2.5 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 3.8 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 3.8 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 1.5 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 30.6 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Marketing Data Reporting via MCP: 5.3 / 6. This was above the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting online sales progress, reporting digital media data, and reporting business outcomes beyond revenue.
- Channel-Level Optimization with MCP: 6.5 / 8. This was above the eight-vendor average of 5.2 / 8. Within this category, its highest-scoring criteria included comparing scenarios with a baseline plan, retrieving saved plans and assumptions, and creating an optimized channel allocation.
- MCP interoperability and workflows: 3.8 / 5. This was above the eight-vendor average of 3.4 / 5. Within this category, its highest-scoring criteria included documenting connections with Claude and ChatGPT, returning reusable structured data, and supporting external agent automation. Analytical Backbone had the same normalized category score.
Where scores were lower
- Enterprise-Grade Platform: 1.5 / 6. This was below the eight-vendor average of 2.3 / 6.
- Campaign & Ad Set-Level Optimization with MCP: 3.8 / 8. Despite being one of Triple Whale’s relatively lower-scoring categories, its score was still above the eight-vendor average of 3.0 / 8.
- Incrementality Testing with MCP: 2.5 / 5. Despite being one of Triple Whale’s relatively lower-scoring categories, its score was still above the eight-vendor average of 2.4 / 5.
4. Lifesight (research score: 29.8 out of 48)

Research overview
Lifesight received an average research score of 29.8 out of 48, the fourth-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Lifesight | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 4.1 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 3.8 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 5.6 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 4.1 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 2.1 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 3.5 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 3.5 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 3.0 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 29.8 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Historical Performance Insights & Causal Explanation via MCP: 3.8 / 5. This was above the eight-vendor average of 3.3 / 5. Within this category, its highest-scoring criteria included reporting incremental ROAS and revenue for digital channels, decomposing base, media and non-media performance drivers, and reporting promotion-driven revenue.
- Channel-Level Optimization with MCP: 5.6 / 8. This was above the eight-vendor average of 5.2 / 8. Within this category, its highest-scoring criteria included creating an optimized channel allocation, forecasting a media plan, and reporting marginal returns and response curves.
- MCP interoperability and workflows: 3.5 / 5. This was above the eight-vendor average of 3.4 / 5. Within this category, its highest-scoring criteria included documenting connections with Claude and ChatGPT, exposing tool and metric documentation, and returning reusable structured data. Analytical Backbone had the same normalized category score.
Where scores were lower
- Incrementality Testing with MCP: 2.1 / 5. This was below the eight-vendor average of 2.4 / 5.
- Enterprise-Grade Platform: 3.0 / 6. Despite being one of Lifesight’s relatively lower-scoring categories, its score was still above the eight-vendor average of 2.3 / 6.
- Campaign & Ad Set-Level Optimization with MCP: 4.1 / 8. Despite being one of Lifesight’s relatively lower-scoring categories, its score was still above the eight-vendor average of 3.0 / 8.
5. Measured (research score: 26.6 out of 48)

Research overview
Measured received an average research score of 26.6 out of 48, the fifth-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Measured | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 4.9 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 3.6 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 4.5 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 2.4 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 3.0 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 3.0 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 3.3 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 2.0 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 26.6 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Marketing Data Reporting via MCP: 4.9 / 6. This was above the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting digital media data, reporting business outcomes beyond revenue, and reporting online sales progress.
- Historical Performance Insights & Causal Explanation via MCP: 3.6 / 5. This was above the eight-vendor average of 3.3 / 5. Within this category, its highest-scoring criteria included reporting incremental ROAS and revenue for digital channels, reporting incremental ROAS and revenue for offline channels, and decomposing base, media and non-media performance drivers.
- Analytical Backbone: 3.3 / 5. This was above the eight-vendor average of 2.9 / 5. Within this category, its highest-scoring criteria included calibrating the underlying MMM with incrementality tests, reporting model validation and modelling KPIs, and providing auditable, editable calibration settings in a self-serve interface.
Where scores were lower
- Campaign & Ad Set-Level Optimization with MCP: 2.4 / 8. This was below the eight-vendor average of 3.0 / 8.
- Enterprise-Grade Platform: 2.0 / 6. This was below the eight-vendor average of 2.3 / 6.
- Channel-Level Optimization with MCP: 4.5 / 8. This was below the eight-vendor average of 5.2 / 8.
6. Fospha (research score: 24.3 out of 48)

Research overview
Fospha received an average research score of 24.3 out of 48, the sixth-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Fospha | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 4.3 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 2.4 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 4.6 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 3.3 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 1.3 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 3.8 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 2.0 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 2.8 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 24.3 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- MCP interoperability and workflows: 3.8 / 5. This was above the eight-vendor average of 3.4 / 5. Within this category, its highest-scoring criteria included documenting connections with Claude and ChatGPT, exposing tool and metric documentation, and returning reusable structured data.
- Marketing Data Reporting via MCP: 4.3 / 6. This was below the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting online sales progress, reporting digital media data, and reporting business outcomes beyond revenue.
- Channel-Level Optimization with MCP: 4.6 / 8. This was below the eight-vendor average of 5.2 / 8. Within this category, its highest-scoring criteria included simulating a user-specified budget change, reporting marginal returns and response curves, and forecasting a media plan.
Where scores were lower
- Incrementality Testing with MCP: 1.3 / 5. This was below the eight-vendor average of 2.4 / 5.
- Analytical Backbone: 2.0 / 5. This was below the eight-vendor average of 2.9 / 5.
- Campaign & Ad Set-Level Optimization with MCP: 3.3 / 8. Despite being one of Fospha’s relatively lower-scoring categories, its score was still above the eight-vendor average of 3.0 / 8.
7. Haus (research score: 18.9 out of 48)

Research overview
Haus received an average research score of 18.9 out of 48, the seventh-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Haus | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 3.1 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 1.9 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 2.9 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 3.1 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 2.6 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 1.3 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 2.3 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 1.8 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 18.9 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Incrementality Testing with MCP: 2.6 / 5. This was above the eight-vendor average of 2.4 / 5. Within this category, its highest-scoring criteria included reporting geo-test results, reporting own-media A/B test results, and prioritizing experiments.
- Marketing Data Reporting via MCP: 3.1 / 6. This was below the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting digital media data, reporting business outcomes beyond revenue, and reporting online sales progress.
- Analytical Backbone: 2.3 / 5. This was below the eight-vendor average of 2.9 / 5. Within this category, its highest-scoring criteria included calibrating the underlying MMM with incrementality tests, providing deterministic answers backed by Bayesian MMM, and reporting model validation and modelling KPIs.
Where scores were lower
- MCP interoperability and workflows: 1.3 / 5. This was below the eight-vendor average of 3.4 / 5.
- Enterprise-Grade Platform: 1.8 / 6. This was below the eight-vendor average of 2.3 / 6.
- Channel-Level Optimization with MCP: 2.9 / 8. This was below the eight-vendor average of 5.2 / 8.
8. Northbeam (research score: 17.6 out of 48)

Research overview
Northbeam received an average research score of 17.6 out of 48, the eighth-highest total in this eight-vendor comparison. Claude and ChatGPT each assessed its public documentation on September 23, 2026, across 48 criteria in eight categories.
Research score by category
| Category | Northbeam | Average research score across vendors | Highest research score among the vendors |
|---|---|---|---|
| 1. Marketing Data Reporting via MCP | 5.0 / 6 | 4.7 / 6 | 6.0 / 6 |
| 2. Historical Performance Insights & Causal Explanation via MCP | 1.6 / 5 | 3.3 / 5 | 5.0 / 5 |
| 3. Channel-Level Optimization with MCP | 2.3 / 8 | 5.2 / 8 | 7.9 / 8 |
| 4. Campaign & Ad Set-Level Optimization with MCP | 1.4 / 8 | 3.0 / 8 | 5.0 / 8 |
| 5. Incrementality Testing with MCP | 1.1 / 5 | 2.4 / 5 | 3.6 / 5 |
| 6. MCP interoperability and workflows | 3.5 / 5 | 3.4 / 5 | 4.3 / 5 |
| 7. Analytical Backbone | 1.5 / 5 | 2.9 / 5 | 3.8 / 5 |
| 8. Enterprise-Grade Platform | 1.3 / 6 | 2.3 / 6 | 4.5 / 6 |
| Total score out of 48 | 17.6 / 48 | 27.2 / 48 | 38.3 / 48 |
Percentage of available points
The scores shown are the average of the evaluations by Claude and ChatGPT using criteria & instructions in this article and public documentation reviewed Sep 2026, not direct product tests. Low or 0 scores may reflect limited public documentation, not that the capability doesn't exist. Bars show the score as a percentage of the category maximum.
Where Claude and ChatGPT assigned higher scores
- Marketing Data Reporting via MCP: 5.0 / 6. This was above the eight-vendor average of 4.7 / 6. Within this category, its highest-scoring criteria included reporting online sales progress, reporting digital media data, and reporting business outcomes beyond revenue.
- MCP interoperability and workflows: 3.5 / 5. This was above the eight-vendor average of 3.4 / 5. Within this category, its highest-scoring criteria included documenting connections with Claude and ChatGPT, returning reusable structured data, and exposing tool and metric documentation.
- Historical Performance Insights & Causal Explanation via MCP: 1.6 / 5. This was below the eight-vendor average of 3.3 / 5. Within this category, its highest-scoring criteria included reporting incremental ROAS and revenue for digital channels, reporting promotion-driven revenue, and reporting incremental ROAS and revenue for offline channels.
Where scores were lower
- Campaign & Ad Set-Level Optimization with MCP: 1.4 / 8. This was below the eight-vendor average of 3.0 / 8.
- Enterprise-Grade Platform: 1.3 / 6. This was below the eight-vendor average of 2.3 / 6.
- Incrementality Testing with MCP: 1.1 / 5. This was below the eight-vendor average of 2.4 / 5.
Frequently Asked Questions
1. Which MCP server received the highest overall score?
Sellforte received 38.3 out of 48. Recast followed at 31.4 and Triple Whale at 30.6.
2. Which vendor scored highest for channel planning?
Recast received 7.9 out of 8, followed by Sellforte at 7.4. Recast’s result includes documented optimizer and forecaster APIs.
3. Which vendor scored highest for incrementality testing?
Recast received 3.6 out of 5 and Measured 3.0. Recast’s evidence includes retrieving imported experiment summaries. Measured received full credit for prioritizing experiments. The category separates result retrieval, prioritization and test design, so a single total does not establish end-to-end test execution.
4. Does connecting an AI assistant make an MMM result causal or more accurate?
The connection gives the assistant access to exposed data and tools. The credibility of the result still depends on the underlying model, experimental evidence, assumptions, validation and appropriate interpretation. The research does not test model accuracy or causal validity in production.
5. How should a team use this comparison to create a shortlist?
Start with the workflow that matters to your team, then inspect the relevant criterion scores and source documents. Ask each vendor to demonstrate that workflow using the intended client and access route. Verify returned metrics, supported granularity, model freshness, permissions and any steps that require the platform UI.
Change Log
| Date | Change |
|---|---|
| September 24, 2026 | This research article was published launched, including evaluation criteria and evaluation of eight vendors. |
Evaluation Dates by Vendor and LLM
Model names and evaluation dates below are recorded in the supplied evaluation log. There are 16 assessments: two per vendor.
| Vendor | LLM | LLM version | Evaluation date |
|---|---|---|---|
| Fospha | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Fospha | Claude | Opus 5.5 | September 23, 2026 |
| Haus | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Haus | Claude | Opus 5.5 | September 23, 2026 |
| Lifesight | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Lifesight | Claude | Opus 5.5 | September 23, 2026 |
| Measured | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Measured | Claude | Opus 5.5 | September 23, 2026 |
| Northbeam | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Northbeam | Claude | Opus 5.5 | September 23, 2026 |
| Recast | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Recast | Claude | Opus 5.5 | September 23, 2026 |
| Sellforte | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Sellforte | Claude | Opus 5.5 | September 23, 2026 |
| Triple Whale | ChatGPT | GPT-6 Astra | September 23, 2026 |
| Triple Whale | Claude | Opus 5.5 | September 23, 2026 |
Limitations and Disclosures
Publisher affiliation. Sellforte designed and publishes this framework and is one of the vendors assessed. Claude and ChatGPT assigned the recorded scores. The article’s interpretation is editorial commentary based on those results.
Evidence scope. The evaluation uses vendor-owned public materials and technical documentation. It is not a hands-on test and does not measure pricing, implementation quality, support quality, performance, statistical accuracy or customer outcomes. Missing evidence is not proof of a missing capability.
Snapshot. The recorded evaluations are dated September 23, 2026. Product access and documentation can change. The scores reproduce the supplied workbook rather than silently incorporating new scores during article preparation.
Corrections. Vendors can submit documentation for correction or re-evaluation to research@sellforte.com.
Recommendation for buyers. Use this comparison as a structured starting point for your own evaluation, not as a final answer. We strongly recommend conducting evaluation calls or demos with vendors to verify fit against your organization's specific requirements.
Further Reading and Resources
Explore related guides, vendor comparisons and research on MCP, AI, marketing measurement and media optimization.
MCP, AI and Agentic MMM
- Introducing Sellforte MCP: Bring Sellforte Into Your AI Workflows
- The Rise of Agentic MMM (Marketing Mix Modeling): How AI Is Transforming Media Optimization
- Webinar: Agentic MMM in Action: The Future of Autonomous Media Planning and Buying in Real Time
- What is the Model Context Protocol (MCP)? — Official documentation
- Budget optimization & Scenario planning with Sellforte MCP (Channel-level)
- Marketing Reporting with Sellforte MCP
Vendor Comparisons and Evaluation Guides
- 6 Best Conversational AI Tools for MMM and Incrementality Testing in 2026: An In-Depth Comparison
- Best Incrementality Testing Tools in 2026: In-depth Vendor Comparison
- Best MMM Solutions for Retail Brands: Top Marketing Mix Modeling Platforms (2026 Q3 update)
- Best MMM Tools for Ecommerce Brands: Top 10 Software for 2026
- How to Choose a Conversational AI Tool for MMM and Incrementality Testing: 48 Evaluation Criteria
MMM and Incrementality Measurement Fundamentals
- What is Marketing Mix Modeling (MMM)? A Complete Guide for Marketers
- What is incrementality testing? A practical guide
- What is Causal Marketing Mix Modeling (MMM)?
- What does "Enterprise-Grade" Mean in Marketing Mix Modeling (MMM)?
- Calibrating MTA with Incrementality: From Attributed ROAS to Incremental ROAS
Experiment Design and MMM Calibration
- Calibrating Marketing Mix Models with Experiments and Attribution data
- How to Integrate Experiments Into an MMM Platform: A Practical Guide
- How should we prioritize incrementality tests across markets and channels?
- How should we weight multiple incrementality tests when calibrating an MMM?
- MMM Calibration Explained: What It Is, Why It Matters, and How to Do It
Media Planning and Campaign Optimization
- ROAS, iROAS, miROAS: Choosing the Right KPI for Optimizing Media Spend
- Advertising response curves: What are they and why do you need them?
- Bid Optimization: How to Calculate the Optimal Bid Values for Your Campaigns Using miROAS
- The Shift in Marketing Mix Modeling: Why Campaign-Level Optimization is Taking Over
- The Five Requirements of Autonomous Media Buying and Optimization: M.A.G.I.C.
Original Research by Sellforte Labs
Author

Lauri Potka is the Chief Operating Officer at Sellforte, with over 15 years of experience in Marketing Mix Modeling, marketing measurement and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world’s largest advertisers on marketing optimization. Follow Lauri on LinkedIn.
You May Also Like
These Related Stories

Best MMM Solutions for Retail Brands: Top Marketing Mix Modeling Platforms (2026 Q3 update)

How to choose a Marketing Mix Modeling solution in Retail: 69 Criteria


