Three hundred eighty-eight people show up to an AI learning day at Gap Inc. Every one of them already has Microsoft Copilot on the laptop in front of them. Access is not the problem in this building.
So the researchers split the day and change one thing: the structure around how people use the tool.
The morning group works however it wants. Pair up, open Copilot, write the strategic plan. Nobody tells them how. The afternoon group gets a protocol, and it’s a good one. Meet on Teams, talk the plan through out loud, hand the transcript to Copilot and let the tool draft from what the two of you actually said.
Then a second test, run on individuals instead of pairs. Same company, same tool, same day. Except this intervention isn’t a rule. It’s a conversation about what the thing is. You’ve been treating this like a search engine. Try treating it like a thought partner. A little guided practice, and that’s the whole treatment.
If you searched “how to use AI for marketing,” you were probably hoping for a list. Tools, prompts, a workflow to copy into Monday morning. That’s a reasonable thing to want, and most of what’s written on the subject will hand it to you. But the Gap experiment tested the two moves marketing teams actually make when they want AI to work, and the results don’t go where you’d bet.
The Good Protocol Made the Work Worse
Pairs assigned the structured protocol scored lower than the pairs left alone. Control pairs averaged 15.63 on the graded document. Protocol pairs averaged 10.68, a gap of nearly five points with a standardized effect size of 0.81 (p < .001).
The number nobody quotes is worse. A lot of the protocol pairs never turned in a document at all. Modeled as its own outcome, treatment pairs had roughly one-eighth the odds of producing anything (OR = 0.12).
A carefully designed workflow, imposed on people, made the work worse and made less of it.
The reframe pointed the other way. Individuals who got the partnership training, the conversation about search engine versus thought partner, had about double the odds of landing a top-tier document (OR = 2.07, p = .022).
Now the part where I slow down, because the researchers are careful about this and I should be too. That top-tier result comes out of an exploratory model, not the headline test, and the primary measure hit a ceiling with 68 percent of documents scoring a perfect twenty. The design has a confound baked in: morning was control, afternoon was treatment, so time of day is tangled up with everything. People dropped out unevenly. The grader was a language model that likes longer documents. Many protocol pairs couldn’t execute the protocol as designed. And it’s a working paper by Microsoft researchers about Microsoft’s own tool, one company, one day.
So it doesn’t prove anything. What it does is better than proof, which is complicate a question in a useful direction. If you’d asked me to bet before I read it, I’d have picked the protocol. I build protocols for a living. That’s the part that stuck.
The Layer Under Your AI Marketing Workflows
Cognitive scientists have had a name for this since Johnson-Laird’s 1983 book Mental Models. A mental model is a working sketch of how a system behaves, built out of experience, and it is stubborn. When you sit down with a tool you don’t interact with the tool. You interact with your picture of the tool, and your picture of yourself sitting next to it.
I’ve started thinking about it as an operating system. Training is an app. Your prompt library is an app. Your AI marketing workflows are an app. Install whatever you want, but if the operating system underneath is wrong, everything you install runs badly. That’s why training the whole team on the tool so often changes nothing you can see six weeks later.
Two Models, One Marketing Team
Watch what happens when the belief is delegation. Somebody opens a chat window and types “write me five subject lines for the spring promo.” Thin input, three seconds of thought, pick the least bad one, ship it. The output is the average of everything ever written about spring promos, because the average is all the input asked for. That person will tell you AI is overrated, and given the model they’re running, they’re right.
Now the same task under amplification. The positioning doc goes in. So does the objection the sales team heard twice last week, the open rates from the last three sends, and the reason December underperformed. Then the draft comes back and gets argued with. Twice. That’s the same discipline as stating a problem somebody could disagree with before you let a tool near it.
Both people are behaving sensibly given what they believe the machine is for. Only one of them is getting anything, which is the whole reason AI makes people faster without making them clearer.
You can usually spot which model somebody holds by watching what they do with a bad first draft. Delegation regenerates. Amplification adds the thing that was missing. Same instinct that decides whether anyone is still checking the work three months in.
Your Defaults Are Teaching the Model
Here’s the half I only learned by finding it inside my own machinery.
Your team’s mental model isn’t only in their heads. It’s written into your defaults.
I run marketing at Auspicious on documented playbooks and scheduled agent tasks. Real system, publishes without me. And the LinkedIn side of it was producing exactly the kind of post I’d tell a client not to make. Which is strange, because my own playbook said in writing that a person’s profile outperforms a company page. I knew that. It was documented.
Then I traced what the machinery was actually capable of producing. The routing tables sent posts to the company page. The derivative sequence made every post a summary of a blog article. The link rules put the link in the body. Stack those three together and the only output the system can generate is a short blurb with a link out, posted by a brand account, which is roughly the worst-performing combination on the platform.
Nothing was broken. Every rule executed correctly. The knowledge sat in one layer and the behavior sat in another, and the behavior kept winning while the right answer sat in a file I’d written myself.
You can’t train your way out of that, and you certainly can’t prompt your way out of it. I had to change the assumption the system was built on: that the writing flows outward from the brand. It doesn’t. It flows out from the person, and the brand plays support.
How to Use AI for Marketing, Starting Monday
Pick someone on your team who touches AI every day. It might be you. Skip the questions you’d normally ask, the ones about training and licenses and whether leadership is supportive. Ask instead what they believe they’re doing when they open that window.
Then go look at your defaults. Not your policy, your defaults. Open the content brief template everybody actually uses. Open the campaign request form. Ask what a reasonable person would conclude from that document about what the machine is for.
If the brief has a line for “topic” and no line for “what we know that our competitors don’t,” you’ve told them it’s a search engine. If the approval flow counts pieces shipped and never asks what got sharper, you’ve told them it’s a content mill. Whatever they conclude from the form, that’s the thing you’re actually running, and it’s teaching louder than any training session.
That’s the layer. Everything else is an app.
Frequently Asked Questions
What’s the fastest way to start using AI for marketing?
Put real proprietary input in front of it before you ask it for output. Your positioning, your customer’s actual objections, your numbers from the last campaign. Most disappointing AI output traces back to a thin prompt, not a weak model. The speed comes from having something worth feeding it.
Do AI marketing workflows and prompt libraries actually help?
They help people who already hold a useful model of the collaboration and they do very little for people who don’t. The Gap field experiment found a well-designed collaborative protocol associated with lower document quality and dramatically lower output than letting pairs work naturally. Structure sits on top. The model sits underneath.
How do I tell whether my team has the wrong mental model?
Watch what happens after a bad first draft. If people regenerate, they’re treating the tool as a vending machine and hoping for better luck. If they add the context that was missing and push back on the draft, they’re treating it as a collaborator. You can also read it off your own templates: whatever your brief makes easy is what your team believes the tool is for.
Should I train my marketing team on AI tools?
Tool training is worth doing and it isn’t the constraint. Access and skill are the layers most programs work on, and they’re the layers most companies have already covered. The constraint is what people believe the collaboration is, plus the defaults in your system that keep teaching them. Fix the belief and the defaults, then the training has somewhere to land.
Working out where AI actually fits in your marketing, and where it’s quietly making things worse? That’s the kind of question fractional marketing leadership exists to answer.
