Most marketing procurement teams can answer parts of the “did we get what we paid for” question. Very few can answer all of it at once, for every agency and every market. Takeaways from the September WMP Fall Series session with Peter Grossman of Flock Associates.
Ask most marketing procurement teams a simple question: across your agency roster, are you buying the right resources, at the right price, against the right scope, and getting the performance you expected? Most of us can answer parts of that question. Very few of us can answer all of it at once, with confidence, for every agency and every market.
It’s not for lack of data. Large advertisers have plenty of agency data. The problem is that it lives in different places, uses different definitions and runs through different processes. Scopes are built and approved in one set of spreadsheets. Rates and fees are negotiated in another. Reconciliation is its own exercise. Performance reviews happen once a year, often in a separate system, and rarely feed back into the commercial relationship.
That fragmentation was the starting point for the September session of the Women in Marketing Procurement (WMP) Fall Series. Peter Grossman, Head of Americas at Flock Associates, walked us through how scope, fee benchmarking and agency performance can be managed as one connected flow instead of three disconnected exercises. The demo was the vehicle but the ideas behind it are what matter for anyone managing an agency ecosystem. Here are the ones I think every procurement leader should take away.
The real problem: three processes, three versions of the truth
Peter summed up the pain points that most people on the call recognized immediately:
- Scopes and appraisals done by hand across a sprawl of Excel files.
- Fee transparency that varies by agency and by market.
- Job titles and naming conventions that make like-for-like comparisons difficult, if not impossible, unless you want to spend your time guessing what each FTE does and then evaluating how it fits across your entire agency ecosystem.
- Performance evaluations disconnected from the commercial relationship.
For procurement, that creates a fundamental problem. When scope, commercial management and performance sit in separate processes and systems, it becomes very hard to answer the “did we get what we paid for” question consistently, especially across a large roster.
The fix is not simply better administration. It is connection. The value comes from being able to move from what you asked the agency to do, to what resources it proposed, to what it cost, to how the agency actually performed, without rebuilding the picture from scratch each time.
Start with a common language for scope
The most underrated step in agency management is standardization. If you manage dozens or hundreds of scopes across multiple agencies, brands and markets, getting everyone onto a consistent taxonomy creates visibility you simply don’t have when every scope lives in its own spreadsheet.
In practice, that means:
- A deliverable library. Define scope deliverables once, at a macro level (“brand strategy,” “big idea development,” “governance and management”) and a micro level beneath them, and reuse those definitions across scopes. Flock’s library holds roughly 1,000 deliverable descriptions, which is a useful starting point if your own library is thin or you are starting from scratch.
- A standard set of job titles. Agencies choose roles from a common list, around 500 titles in Flock’s case, instead of inventing their own. This is what makes true apples-to-apples comparisons possible later.
- Room for every pricing model. Time and materials, asset-based pricing, commissions and pass-through costs should all fit in the same structure. Remuneration models are evolving quickly, and your scope framework should not force every relationship into one model.
One practical point came up in the Q&A: who builds the scope? The client usually starts it (through a pitch, brief or simple conversations) and the agency responds, often suggesting additions the client had not considered. Some clients hand the initial upload to their advisor to save time. One participant also noted that “most” AI tools can now pull scope details from a contract into a template quite accurately, which is a sensible way to take the grunt work out of the first draft.
Ask agencies to show their workings
The next step is how agencies respond. Instead of a single fee number, the agency provides the building blocks: where the work will be delivered from, its overhead and markup rates, annual hours, any agreed bonus and malus percentages, the roles it will use with an hourly rate for each, and the hours per role against each deliverable.
This matters because delivery location is a real cost lever. A scope may be delivered to the US, but the agency may choose to produce some of the work in Canada, Colombia or Malaysia to drive efficiency. You want that visible, and you definitely do not want it buried in a blended rate.
Once responses come back, especially in a pitch where every agency is pricing the same scope, the comparison options open up. You can compare agencies side by side on total cost, hours, FTEs and blended hourly rate. You can then drill into cost by location, by discipline, by seniority band, by individual role and by deliverable, all against benchmark.
The goal is not just finding savings. As Peter put it, it is about right-sizing the structure. Sometimes the data shows an expensive role that deserves a conversation. Sometimes it shows the opposite: every other agency put senior people on a deliverable and this one did not, which raises the question of how the work will get done well. Both findings strengthen the probability of successful outcomes for all parties.
Benchmark like for like, and keep doing it
Benchmarking is only as good as its comparability. The benchmark should match the market, the role and the deliverable: a scope delivered from Colombia should not be benchmarked against US rates. Flock benchmarks by market, and its database contains roughly $3 billion of negotiated agency fees (Flock Associates). Two details stood out to me:
- The data stays fresh. Nothing older than 18 months stays in the database.
- Some of it is forward-looking. It includes agency pricing submitted in pitches for work in future periods.
Flock reports average savings of around 30% when clients run a full benchmarking exercise.
The bigger shift is in how often we benchmark. Too many organizations treat fee benchmarking as a once-every-three-years event, tied to a pitch or contract renewal. In reality, rates and resourcing change continuously: new markets, new roles, new headcount requests. Benchmarking should be part of ongoing relationship management, not a one-off.
That is where quick rate-card checks are useful. Flock recently launched an Agency Rate Check tool: enter an agency’s overhead, markup, annual hours and role rates, and it shows where each element sits against benchmark. In the demo, a sample rate card came back with overhead well below benchmark, markup well above it and annual hours within the expected range. That is exactly the level of detail you need to negotiate.
When agencies won’t share the breakdown
One of the best questions of the session was also the most realistic. Outside an RFP, agencies often won’t disclose overhead and markup. What then?
You can still benchmark the fully loaded hourly rate against the market. That tells you whether a rate is high or low. It does not tell you why. Without the component rates, you can say “these rates look high, please address it,” but you can’t point to margin or overhead as the driver. Transparency on the building blocks is what turns a rate conversation into a fact-based discussion.
If you do not have access to the true numbers, you can still add estimates into your rate verification and through that better understand what it is that your agency is not willing to share with you. Sometimes that can be a more insightful discovery than having all the numbers at your fingertips.
On annual hours, in my experience most agencies in the US now work to about 1,800 hours a year, sometimes 1,900. If an agency’s number sits well outside that range, it deserves a question.
Benchmarks also need a qualitative layer. National benchmark data does not break out individual cities, so a lower-cost hub such as Austin, Texas, may make an agency’s rates look favorable against a US benchmark when they are merely in line locally. Good advisors add that context instead of letting the number speak alone.
Agency performance: connect it to the commercial relationship
The second half of the session covered agency performance evaluation. Most organizations already run some version of an agency review. What differs is design and follow-through.
A few principles I would highlight:
- Choose the right lens. A 90-degree review is client on agency. A 180-degree review adds agency on client. A 360-degree review adds either an agency self-assessment or agencies evaluating each other. A participant who had run 360s made a strong case for them: they surface blind spots, because it is always easier to say the other side needs to change. Agency-on-agency reviews are rarer, since many clients worry about politics and objectivity, but they can be revealing.
- Keep it short enough to be answered well. For a full year-end review, 20 to 25 questions are the practical limit before fatigue sets in. Many clients add a lighter pulse check of about 10 questions quarterly or twice a year.
- Use benchmarkable questions. Standard questions that have been used across many evaluations let you compare scores against a wider benchmark. Custom questions are fine but usually lose that comparison.
- Require the “why.” On a five-point scale, any score other than a neutral three should prompt a written comment. The qualitative answers, sorted into strengths and areas for improvement, are often where the action plan comes from.
- Weight where it matters. Results can be weighted by seniority, market, category or question, for example giving lead markets more weight in a global review.
- Compare agencies carefully. Scores for several agencies in the same capability area can be compared side by side, but the client should decide whether agencies ever see that comparison.
Flock reports performance improvements of 15% to 20% among clients using its evaluation approach (Flock Associates). Whatever the tool, the principle holds: evaluation drives improvement when it is consistent, repeated and tied to the commercial relationship, not filed away after the annual meeting.
What this looks like in practice
I have had the pleasure of working on several Flock projects, and one engagement shows how these pieces come together over time.
The client is a global consumer goods company. The client had just completed a major renegotiation with its media agency: not a pitch, but a new MSA that added markets and introduced a performance-related incentive. Flock came in at the point where the real management work begins.
The client started with the agency evaluation. The same questions run across every market, with a baseline review for the newly added markets, then mid-year and year-end reviews, moving to quarterly next year.
Next came the incentive. Across multiple markets, calculating a performance-related incentive gets complicated quickly. This incentive is based on three inputs: the evaluation scores, media savings verified by an independent media auditor, and innovation. Flock built a calculator that takes those inputs and produces the payout automatically. Having worked on both the agency and the client side, I have built those spreadsheets myself. Every time, the agency builds its own model, the client builds another, and the two end up a million dollars apart. An agreed, shared model settles the number at period end and lets everyone sign off.
The rate benchmark came later, almost by accident, when we realized it was available. We now use it live as markets are added. When the agency proposes another person to be added to the account, we check the rate against benchmark before agreeing. The client had previously paid an annual license for a competing benchmark database that it checked about once a year, and those on-demand checks cost far less. The next step is the creative agency.
That progression, from appraisal to incentive to rate benchmarking, is what an integrated approach often looks like in practice. You don’t need to adopt everything at once. You start where the pain is greatest and build out as the relationship matures.
The takeaway for procurement
The question in the session title, “did you get what you paid for?”, should not be answerable only once every few years, after a lengthy audit. It should be answerable at any point in the relationship. That requires four things:
- A common language for scopes, roles and deliverables.
- Transparent pricing broken into its building blocks.
- Continuous, like-for-like benchmarking.
- Performance evaluation tied directly to the commercial terms.
Whether you get there with a platform, an advisor or a disciplined internal process, the goal is the same: stop treating scope, fees and performance as separate exercises and start managing them as one relationship.
Thank you to Peter Grossman of Flock Associates and to everyone who joined and asked such sharp questions. The WMP Fall Series continues this autumn.
Want to know whether you are getting what you pay for? We help brand-side teams connect scope, fees and agency performance. Get in touch for an initial conversation.
Sources
Figures are Flock-reported data, presented by Peter Grossman at the WMP Fall Series session “Did You Get What You Paid For With Your Agency?” on September 22, 2026.