TMOD LogoTMOD

Website content audit tool

15 content checks with an AI reviewer pass, depth, originality, thin pages, E-E-A-T and scaled-content risk.

What TMOD checks

  • Word counts per page against a 300-word minimum and a 600-word target, with a separate thin-page pass flagging anything under 150 words and anything under 100 as very thin.
  • Near-duplicate content across real content pages at a 0.8 similarity threshold, ignoring archives and pagination.
  • Total published pages against a 15-page minimum and a 30-page target, judged by sitemap and link discovery rather than only the pages we fetched.
  • Scaled-content risk: an AI originality read, corroborated by a structural-uniformity heuristic that flags six or more pages with a word-count coefficient of variation under 0.12, the statistical fingerprint of templated mass production.
  • Presence of About, Contact, Privacy and Terms pages, detected in any language by an AI classifier rather than English URL patterns.
  • Author bylines and E-E-A-T trust signals, readability, image alt text coverage, publishing cadence, content freshness, and the language the site is written in.

Every check, explained

15 checks run in this audit. 5 of them have a page of their own with the exact threshold and how to fix it.

Why it matters

Content problems are the most common reason a site is judged low quality, and the hardest to see from the inside. You know what every page is for; a reviewer arriving cold sees a directory of near-identical 200-word pages and reaches a different conclusion in about ten seconds.

The uniformity heuristic is worth understanding because it catches a specific failure that no word-count check will. Twenty pages that are each 480 words long is not a natural distribution, real writing varies. A coefficient of variation under 0.12 across six or more pages is a strong signal the pages came out of a template or a generator, and it is exactly what the scaled-content-abuse policy targets.

This preset includes the AI reviewer pass, so it reads your content in any language and judges it on its merits rather than on keyword heuristics. That makes it slower and more expensive than the technical audit, and considerably more useful when the question is 'is this good enough'.

How to fix it

01Consolidate before you create

If the report flags thin pages and near-duplicates together, you almost certainly have several pages competing to cover one topic. Merge them into a single thorough page, redirect the old URLs, and you fix both problems and the internal-link dilution at the same time.

02Break the template

If uniformity is flagged, the fix is not to add words. Vary the actual structure: some pages should be short and direct, some long and thorough, with different section counts and different kinds of evidence. Content that was worth writing individually looks individually written.

03Add the person behind the writing

E-E-A-T signals are concrete, not mystical: a named author on each article, a real bio explaining why that person knows the subject, and a link between the two. On a site with no author attribution anywhere, adding it is one of the highest-leverage content changes available.

04Write the essential pages properly

About, Contact and Privacy are checked for existence, but a two-line About page is nearly as bad as none. The About page is where a reviewer decides whether a real person or organisation is behind the site, it deserves a few real paragraphs.

What low value content means in practice

The phrase is doing a lot of work in Google's vocabulary, and it does not mean badly written. It means a page that does not repay the visit: the answer is already in the title, the body restates the question three times, and the reader leaves knowing exactly what they knew before. Length is only how that usually shows up in a word count.

The recognisable shapes are worth naming, because most sites have one of them rather than a general quality problem. Location and service pages generated from a list, differing by a place name. Product or comparison pages assembled from a feed with no first-hand use of anything. Tag and category archives with one item in them. Round-ups that summarise five articles you did not read. Definition pages that exist to catch a keyword variant of a page you already have.

What they share is that nothing on the page could only have come from you. That is the practical test, and it is a better one than any threshold here: if a competitor could produce the identical page from the same public inputs in ten minutes, the page is not carrying its own weight, whatever the word count says.

Turning the report into a content plan

The number that moves a judgement is a ratio, not a total. Forty good pages next to sixty stubs reads worse than forty good pages on their own, which means deleting is a legitimate and fast way to improve a site. Start there: pages you would not send a reader to should be merged, redirected or removed, and empty archives should stop being indexed.

Consolidation comes next, and it is where most of the recoverable value sits. Every group of pages circling one topic becomes a single page that covers it properly, with the old URLs redirected so nothing that already points at them is lost. This is also the fix that breaks a uniform structure, since a merged page is a different length and shape from the ones it replaced.

Only then is publishing more worth the effort. Cadence matters more than volume here, since a steady stream reads as a live site while a burst of twenty pages in one afternoon reads as exactly what it is. Re-run the audit after each batch, and check that the number of flagged pages is falling rather than the total number of pages rising, then confirm the rest of the picture with the full audit before you act on it.

Questions

Does this detect AI-written content?

It flags scaled-content risk, which is not the same thing. Google's policy does not prohibit AI assistance; it prohibits mass-produced pages made for search rather than people. We look for the signals of that, near-identical structure, uniform lengths, low originality, no editorial voice. Carefully edited AI-assisted writing that varies naturally will not trip it, and neither will genuinely human writing that happens to be templated.

Why is a short page flagged when it is meant to be short?

Essential and utility pages, contact, privacy, terms, are identified by the AI classifier and excluded from the thin-page check, in any language. If a genuinely short page is still being flagged, it is being read as a content page. That is usually a signal worth taking seriously: if the classifier cannot tell it apart from an article, a crawler will not either.

What counts as duplicate content?

Two real content pages whose extracted text is 80% or more similar. Archive pages, category listings and paginated views are excluded because they overlap by design and Google does not treat them as duplicates. A flag here usually means either the same article published at two URLs, or several pages built from one template with only a name swapped.

How recent does content need to be?

There is no fixed cutoff, because it depends entirely on the subject. The freshness check looks at your publishing pattern and the age distribution of your pages relative to each other. A site whose newest post is three years old reads as abandoned regardless of topic; a site with a steady cadence is fine even if individual pages are old.