The world of AI slides is changing fast and within days of each other in September 2026 the two largest AI labs, Anthropic and OpenAI, released new flagship models that are – among other things – reportedly much better at creating slides. Anthropic launched Claude Fable 5.1 on September 1, and OpenAI began rolling out GPT-6 Astra shortly after. OpenAI describes Astra as its best model for producing slides that are well laid out and "convey key points with a structured narrative".
In our previous article, we tested Fable 5.1 on three typical consulting tasks. This time, we decided to pitch the models against each other to see which outputted the best consulting-level slides. We compared the time each model took, what it cost and, most importantly, how close the output came to a McKinsey or BCG deck. We then asked both models to make their decks more consulting-like and scored the results again.
How we tested Fable 5.1 and GPT-6 Astra
We used a real, publicly available industry report: Oil & Gas in Egypt: Market Report, a 52-page report written by Osama Kamal for Norwegian Energy Partners (NORWEP) in August 2019. The report covers Egypt's economy, investment rules, drilling activity, the main operators, production sharing contracts, the Zohr and West Nile Delta gas projects, and a long newsfeed of deals.
It's a typical input for a market study: rich in data, but with no storyline and no recommendation.
We ran the same two prompts in each model, both on the High effort setting:
Round 1: Turn this report into a deck
Round 2: Can you make it more consulting-like so it looks like decks that McKinsey or BCG consultants make in terms of level of detail and complexity of slides?
We scored every deck against the same nine criteria we use for all our AI slide tests: layout, focus, action titles, tone, visual balance, storyline, coherence, client-readiness and overall MBB look and feel:
- Does the layout of the slide effectively communicate the key message? A slide's layout should emphasize the data on the slide. It makes a big difference if you use a table vs. a chart vs. a diagram to get a key message across and make it easily understandable. The art of choosing layouts is trained within consulting houses through years of seeing examples and getting feedback.
- Does each slide manage to only keep information on it that’s relevant for the main message of the slide? Just like it's important to choose a layout that communicates the key message, it's also important to clean the slide for any details or info that make it more difficult to grasp the key message. It's fine adding detailed info (in fact, it's encouraged with techniques like the pyramid principle) but having data on the slide that is not directly connected to the main message makes it confusing.
- Do the slides have good action titles? Action titles summarize the main message of the slide and are a crucial part of any client deliverables.
- Is the level and tone of text on the slide consulting-like? Consulting slides are often quite detailed and with a tone that carries a certain point of view or conclusion on whatever the slide says.
- Is the entire deck visually balanced? This is one of the most important aspects of consulting decks. The balance between layouts needs to be interesting and support the arc of the storyline. So a deck with three tables in a row will fail, while a deck that has a chart, a diagram, a table etc. will not.
- Does the deck follow a structure that mirrors how a consultant would structure it? Storylining is a key part of what makes consulting decks good. Different types of decks should follow different storylines, just like the audience and setting will impact how the deck is structured.
- Are the messages and numbers coherent throughout? This is a given.
- Is the output client-ready? Would we send this deck to a client?
- Do the slides look and feel like McKinsey or BCG consulting slides overall? Not including overall design choices like colors and fonts, do the slides look and feel like real consulting slides?
We rated each criterion Yes, Partially or No.
Round 1: Can Fable 5.1 and GPT-6 Astra turn a PDF report into a PowerPoint deck?
For our first round, we wanted to test a simple prompt to get a baseline of the output. For both models, our prompt was simply "Turn this report into a deck."
Claude Fable 5.1: An okay market brief, but not McKinsey-level
Our first test was with Claude Fable 5.1. We uploaded the PDF report and asked Fable to turn it into a deck.
Before building, Fable 5.1 asked two questions: how detailed the deck should be (it recommended around 15 slides) and who the main audience was. It recommended framing the deck for suppliers and service companies, "as the NORWEP report intends". We accepted both recommendations.
Fable 5.1 returned a 16-slide market briefing for suppliers with an executive summary, a body of slides detailing various analyses from the PDF report, and ending with a suggested focus for suppliers based on its own interpretation.
Judging the deck against our criteria:
All in all, a solid first draft but nowhere near the level of slides produced by houses like McKinsey and BCG.
GPT-6 Astra: A basic, text-heavy deck
Before building, GPT-6 Astra asked which style the presentation should use. It offered to use an uploaded template or one of its own designs, such as a market trends report or a business review. We chose the blue-and-white Business Review.
GPT-6 Astra also produced 16 slides that follow the structure of the report: economy, rigs, drilling, production, infrastructure, megaprojects, gas targets, licensing, the players, bid rounds, services and downstream projects.
The output looks pretty basic, and rating the deck against our consulting criteria it falls short:
The raw outputs of the two models were fine as a basic summary, but far from an actual consulting deck. Let’s now see if we can move it more towards that.
Round 2: Can Fable 5.1 and GPT-6 Astra make McKinsey-style slides?
In the same chats as the previous decks were made, we now asked both models the same follow-up question to push them towards consulting standards:
Claude Fable 5.1: Closest to a real MBB deliverable
The second prompt changed Fable 5.1's output completely. It rebuilt the deck as a 23-slide detailed deck with a contents page, numbered sections, section trackers, more in-depth recommendations and implications for the “client”, and much more detailed charts and layouts.
Like before, we measured the deck against our consulting criteria:
This deck is the closest to a real consulting deck and could be lifted with few, but effective tools.
GPT-6 Astra: Slightly better, but nowhere close to a real consulting deck
In our final test, we asked GPT-6 Astra to make the deck more consulting-level. It kept the 16-slide structure it had before, but rebuilt most of the slides.
Unfortunately, although the visual style became tighter and closer to a consulting look, it still was nowhere close to the McKinsey benchmark:
The second-round deck from GPT-Astra was very disappointing. It seemed to focus on formatting and action titles, but didn’t change the meat of the slides in terms of layout, density etc.
Scorecard: Rating Fable 5.1 vs GPT-6 Astra on 9 consulting criteria
The table below summarises how all four decks scored against our nine criteria. A Yes means the deck met the standard of a McKinsey or BCG deck on that criterion while a Partially means it got some of the way there.
Based on out test, three findings stand out;
- Fable 5.1 is much better at the visuals than GPT-6 Astra:
Fable 5.1 chose good and effective layouts from the first prompt e.g., an org-style diagram of who the buyers are, an indicator table with up and down arrows, and charts that match the data. GPT-6 Astra relied on text-heavy tables and boxes in both rounds, so its layout and visual balance scored No each time.
- The second prompt helped Fable 5.1 far more than GPT-6 Astra:
Fable 5.1 moved from generic deck to something much closer to a real MBB deliverable based on the simple prompt. It rebuilt the entire deck.
GPT-6 Astra seemed to add some details and adjust text, but did not materially change the deck.
- Fable 5.1's best deck still needs a consultant:
It's the closest of the four decks to a real consulting deliverable, but it contains made-up content that has to be checked line by line, old-school takeaway boxes and a "so what" label in the executive summary. It also misses the details that make real MBB slides work: call-outs and annotations on charts, and titles that draw out the implication rather than stating the fact.
Time and tokens: How long Fable 5.1 and GPT-6 Astra take to build a deck
We also compared how much it cost and how long it took for all four runs:
Fable 5.1 wrote more, roughly 126,000 output tokens across both rounds against about 57,000 for GPT-6 Astra, which fits its larger, more detailed second deck. GPT-6 Astra processed more in total, about 11.5 million tokens against 8.5 million, but almost all of it (about 11.1 million) was cached input.
The two apps count usage slightly differently, so treat the figures as indicative rather than an exact cost comparison. In practice, both models built a full deck in minutes, and the time that matters is the review afterwards.
Conclusion
Neither Claude Fable 5.1 nor GPT-6 Astra can build consulting-grade slides from a simple prompt. Both returned fine decks, but none that were client-ready or close to real McKinsey or BCG decks.
That being said, Fable 5.1 is the better choice for consulting slides. Its first deck was already more visual than GPT-6 Astra's, and when we asked for consulting-level detail it delivered proper action titles, a consulting tone, and a clear structure. GPT-6 Astra improved its titles but stayed text-heavy and kept the same structure.
However, even Fable 5.1's best deck is a strong first draft, not a finished deliverable. To improve the output, ask explicitly for consulting-level detail, give the model good reference slides, and build on your own slide master. See our tips for getting better decks out of Fable 5.1 in our related article.