Wissen

Study: AI Texts Are Only as Good as Their Prompt Guardrails

Content Pipelines

On the screen lie two texts side by side, both from the same assignment, both produced by the same AI within minutes of each other. The text on the left reads like an official notice: noun chains, passive sentences, not a human in sight. The right one sounds like a children’s book: eight words per sentence, every bit of expertise watered down. Between the two lies exactly one variable, the instruction that accompanied the assignment.

As pointed as this scene is, a German corpus analysis has now measured exactly this pattern across 2,112 AI texts. The result: whoever wants to improve AI texts finds the most effective lever, according to the data, in the guardrails the prompt sets for the model. And in the “human” editing that follows. A note on scope up front: this is a German-language study of German-language texts, and my own measurement run below uses a German readability formula. I carry the method and the lesson; the figures stay citable with their origin attached, as German-language findings rather than general English-language results. What the numbers say, where their limits lie, and what that means for your content production follows below.

The German study by WORTLIGA and SISTRIX

The Munich text-analysis provider WORTLIGA, together with the SEO tool maker SISTRIX, had 2,112 B2B texts generated between May 11 and June 11, 2026, by their own account, using three large language models (claude-opus-4-7, gemini-3.1-pro-preview, gpt-5.5), spread across eleven text genres from the LinkedIn ad to the specialist article, eight industries, and eight deliberately different prompt styles.

Measurement used the in-house WORTLIGA score from 0 to 100, which, besides readability, factors in passive constructions, nominal style, filler words, and clichés. Two things about this measure need to be said openly. First, the score measures readability and comprehensibility. About effectiveness, conversion, or leads it says nothing, and about factual correctness just as little. The study title “How effectively does AI communicate in marketing and sales?” therefore promises a bit more than the method measures.

As a target range, the authors name 60 to 70 points out of 100, depending on audience and topic. Values above 80 they themselves rate as a sign of oversimplification. To place the scale: the three model averages lie, per the study, between 37.7 and 47.7 points, on average 44.1. No AI model reaches the target range on average.

A bureaucratic prompt yields officialese, the demand for more clarity yields children’s books

The core finding fits into two pairs of numbers. Switch the prompt style, and the category averages wander across all models from 4.4 to 79.4 points, once across almost the entire scale. Switch the model, and the averages move from 37.7 to 47.7. Ten points. So the instruction shifts the result across a whole field, while the model choice barely leaves the center circle. An honest addition belongs with it: which factor explains how much of the spread, the study does not show. The visual impression is clear nonetheless, and the authors have given it a name: the chameleon effect. The models mirror the style of the instruction you give them.

What does that look like? A deliberately bureaucratic prompt pushed the category average of all three models down to 4.4 points. Officialese on order. The instruction to write extremely clearly catapulted the average up to 79.4. Sounds like a win, but per the study it reads like a children’s book: staccato sentences, language level A2, useless for expert readers. So is the swinging readability down to the model? The data give the instruction the far larger role.

[Infographic placeholder: span of the WORTLIGA readability averages — 75 points across the eight prompt styles (4.4 to 79.4), ten points across the three models (37.7 to 47.7). Source: WORTLIGA × SISTRIX, German-language study, 2026. EN adaptation follows separately.]

The problem: 52 percent of marketing departments lack the skills to steer AI models, says Bitkom

That the instruction decides the readability of AI texts would be a footnote if companies had this steering firmly in hand. A status report by the German digital association Bitkom paints a different picture: 51 percent of the companies surveyed agree that generative AI already handles large parts of the creative work in marketing. 52 percent at the same time agree that marketing departments often lack the skills to use AI applications sensibly.

It follows that between these two assessments lies exactly the gap this article is about: the tools are in use, but professional steering lags behind. For context: 180 companies from the Bitkom digital-industry network were surveyed (calendar weeks 44 to 50 of 2025); the results are, by their own account, not representative. A snapshot of sentiment, not a market measurement. As such, it matches what many content leads experience daily.

Improving AI texts takes more than the best prompt

The study also tests the reverse: what does a cleanly built prompt achieve? The test’s best-practice style, with a named target audience, the benefit in the first sentence, and a strict ban on passive and clichés, delivered the best category average at 49.3 points. For two of the three models, passive errors dropped to exactly zero. Guardrails in the prompt work, measurably and immediately.

And yet this value, too, stays under the study’s 60-point threshold. The documented top values of the test come from the extremely-clear category, at the described price of children’s language. A prompt that lifts both over the threshold, combining clarity with expertise, the study does not show. Editorially revised texts it did not test at all. The consequence we draw matches the authors’ main recommendation, a “systematic quality assurance”: after the best prompt, editing begins. Check, cut, sharpen, hold the expertise where the score tempts toward simplification. That is our conclusion from the data.

The self-experiment: our posts in our own measurement run

Whoever cites such numbers should not shy from testing on their own material. So I ran our already-published Wissen posts through a measurement run of my own (measurement date: July 14, 2026). I measured with the German Flesch variant after Amstad, which computes a readability value from 0 to 100 out of mean sentence length and syllables per word. Why not use the WORTLIGA score directly? For a good reason: the Flesch formula is open; anyone can recompute it with three lines of code or, if need be, with pen and paper. One caveat belongs with it: this Amstad variant is calibrated for German, so the figures describe German-language texts and do not transfer to English readability as-is.

The result: our posts lie between 46.4 and 55.8 points, on average 50.1, at a mean sentence length of 13.4 words. The narrow span is the real message to me: a consistent language level across all posts. Which share of that comes from the guardrails in the prompt and which from the editorial check, this simple measurement does not separate. It shows the result of the whole process, no more. How reviewed, consistent content also lowers the risk that an AI explains your offer wrongly (in German) we have described elsewhere.

My recommendation:

Write your guardrails down cleanly once: active-voice requirement, cliché ban, a sentence-length corridor, the benefit for the target audience in the first sentence. And still treat every result as raw material. The prompt lifts your AI texts to a reliable level. Whether a text carries your offer is decided by review from people who know the topic and audience.

Two texts, one variable

Back to the two texts from the start. Official notice and children’s book come from the same machine, the same assignment, the same model. Which of the two becomes more likely, you decide. In the instruction, before the model writes its first word. Whether a text that carries your offer results from it is decided afterward, at the editor’s desk. AI texts are only as good as their guardrails. And guardrails are editorial craft, in the prompt as much as in the review that follows.

Sources: WORTLIGA × SISTRIX, “Wie wirksam kommuniziert KI in Marketing und Vertrieb?” — analysis of 2,112 AI texts, 2026 (in German) · SISTRIX, “KI-Texte im Lesbarkeits-Check” (chameleon effect), 2026 (in German) · Bitkom marketing status report, 2026 (in German) · PR DESK own Flesch-Amstad measurement run, July 14, 2026 (internal).

This English article was adapted with AI support in an editorially reviewed production chain, with a human holding every release. Adapted from the German original: https://prdesk.de/wissen/ki-texte-verbessern/

Aus Ihrem Fachwissen wird ein Content-System. Lassen Sie uns klären, welches Thema Ihre erste Content Pipeline trägt. Erstgespräch vereinbaren →