Don't Use an Agent Like a Chatbot
Many people use agents much as they use chatbots. Take a data task: someone on the business side asks a question, the agent returns its findings, and the conversation ends there. The answer may be richer than what traditional BI provides, but people still have to build the audience, develop a strategy, and set up experiments.
A while ago, in “Can Coding Agents Do Data Mining?”, I wrote about using a coding agent to investigate subscriber churn.
We began by looking closely at a few high-value users and quickly reached a conclusion: they preferred educational content and had left because the supply of that content had not kept up. After the agent expanded the sample and extended the observation period, that conclusion fell apart. Users interested in educational content were only a small minority. Many subscribers continued using the product after their subscriptions expired; they had simply stopped paying. For many who did leave, content consumption had already been declining for some time.
That article stopped at the findings. In the actual project, we handed the follow-up work to our own data agent. It integrates data retrieval, user profiles, behavior analysis, audience building, and experiment analysis. As new evidence emerges, it adjusts the analysis and strategy, eventually producing audiences and plans ready for online testing.
This work used to pass back and forth among analysts, operations staff, data engineers, and product managers. Now it happens within a single task.
After the Analysis, Build Audiences and Experiments
Once the analysis was complete, the agent turned its findings into actionable audiences. It divided users into three lifecycle groups: those still consuming content without renewing, those whose consumption was declining, and those who had been dormant for a long time. It then used profiles, raw behavior data, and audience-building tools to translate the analytical judgments into structured conditions.
With candidate audiences in place, the agent used stratified sampling to check whether the selected users matched the original judgments. Long-term preferences helped identify what content users might need. Recent raw behavior confirmed whether that need was still present. Structured fields controlled subscription status, last active time, and eligibility for outreach.
The actual audience results changed the strategy again. The educational-content audience we had initially envisioned was precise but small. Audiences identified by lifecycle and consumption status were much larger and more worth prioritizing. The business question shifted from “How do we bring back users who have already left?” to “How do we intervene earlier as users begin to churn?”
Next, the agent looked up past experiments on similar audiences. Some broad outreach campaigns had failed to deliver the expected returns and had also hurt long-term value. That evidence ruled out simply sending more ads or offers. The strategy shifted toward bringing users back through content, addressing non-renewal, and tighter frequency controls.
For users who were still consuming content without renewing, the plan was to remind them of subscription benefits when a content need arose, with copy based on the topics and specific content they had been consuming consistently. For users whose consumption was declining, the priority was helping them continue with relevant content, with less reliance on direct price incentives. Evidence was weakest for long-dormant users, so they received lower priority and tighter outreach cost controls.
The agent ultimately prepared several sets of executable audience conditions, along with content strategies, triggers, frequency caps, exit conditions, primary metrics, and guardrail metrics. It could create the experiment configurations directly, leaving the user to give final confirmation in the A/B testing system.
Analyze causes → Identify audiences → Develop strategies → Create experiments → Launch → Read results
This workflow includes human confirmation, but no manual handoffs. Business users do not have to download lists, copy audience conditions, or explain the findings again to the next team.
The workflow is still offline. The agent analyzes the problem, prepares strategies, and creates experiments. It does not make real-time decisions for each production request. Once the user confirms, existing recommendation models and business systems execute the strategies.
What Changes When the Steps Are Connected
None of the individual steps is new. Data platforms can query data, profile platforms can inspect users, audience platforms can create segments, and experiment platforms can evaluate results. In the past, each system handled its own part, while people organized the work in between.
When an agent connects these capabilities, new findings can directly change what happens next. If the educational-content audience is too small, the agent keeps looking for a larger audience worth intervening with, without waiting for someone to submit another request. If past experiments show that broad outreach is risky, it adjusts the strategy and constraints before taking a flawed conclusion into production.
The most immediate value is avoiding wasted effort. Had we acted directly on the initial small-sample conclusion, resources would have gone toward a tiny audience, and the outreach might have repeated past failures.
The same analysis can also support multiple experiments. The agent combines lifecycle stages, content needs, and triggers into different plans, rather than putting everyone into one large audience under a single strategy.
After launch, the original task reads the results on a predefined observation schedule, checking consumption, payments, and guardrail metrics together. Findings that hold up inform the next round of strategies. Failed attempts leave a record of where an approach does not work. The analysis remains in the task for later strategies to use.
Turn Data Platforms into Working Interfaces for Agents
Traditional data platforms divide capabilities by product. Queries belong to the analytics platform, user understanding to the profile platform, audiences to the segmentation platform, and validation to the experiment platform. Task state and business judgments remain with people. Each time the work moves to another system, someone has to explain the goal and context again.
Handing this work to an agent does not require rebuilding those platforms from scratch, but it does require changing how they expose their capabilities.
| Existing product capability | Working interface an agent needs |
|---|---|
| Enter queries and view results in an analytics platform | Accept a clearly defined query task and return structured data and actionable errors |
| Inspect users one by one in a profile platform | Read profiles, memories, and raw behavior for specified users and time ranges |
| Configure conditions manually in an audience platform | Convert natural-language judgments into conditions, return audience sizes and samples, and create audiences |
| Search for and configure experiments in an experiment platform | Query past experiments, create experiment configurations, and execute them after user confirmation |
| Log in to and operate each platform separately | Inherit the initiating user’s permissions and record calls throughout the task |
Interfaces alone are not enough. The agent also needs to know which data to query, how metrics are defined, whether results are abnormal, and what past experiments tell us. Table schemas, metric definitions, transformation logic, domain knowledge, and correction records need to become part of the six layers of Data Context it can read. Otherwise, it may complete the entire workflow using the wrong metric definition.
Each interface also needs stable inputs and outputs. Data retrieval, profiles, behavior queries, stratified sampling, audience building, and experiment analysis are exposed as tools. Guidance on how to use those tools for business tasks is organized into Skills. The data team begins building capabilities that can be called repeatedly: the shift from writing SQL to training agents.
Permissions must also extend beyond account controls within individual platforms to cover the entire task. Read-only queries can run automatically. Actions that change external state, such as creating audiences and experiments, must be recorded and confirmed by the user before execution.
Once an experiment is live, the platform also needs to retain the original hypothesis, audience conditions, and strategy version, then return the results to the original task. Validated findings enter subsequent Context, while stable procedures become Skills that inform the next round of planning.
How to Tell Whether an Agent Is Actually Doing the Work
A chatbot delivers an answer, so evaluation focuses mainly on whether that answer is correct. For an agent, the answer is an intermediate result. We also need to ask how far it has advanced the task.
Start with task completion rate. A business question that ends in an analysis report and one that produces actual audiences, experiment configurations, and traceable results represent very different levels of completion.
Then count manual handoffs. User confirmation of permissions or an experiment launch is a necessary control point. But if people still have to download data, rewrite audience conditions, or explain the findings again, the work has not yet been handed over to the agent.
Also measure the time from question to experiment. An accurate analysis that takes weeks to reach the business still has limited value. A useful insight needs to reach validation sooner.
Finally, look at incremental impact in production. Task completion rate and execution speed show whether the agent can carry the work forward. Online experiments tell us whether it creates business value. Those results must also feed into the next round of strategies, or the work starts from scratch again next time.