Brain+AI

SEO and AI visibility 8 min

Programmatic SEO: when it works, when it hurts

Programmatic SEO generates pages in bulk from data, such as one page per product or per city. It works when every page carries data the visitor will not find elsewhere and when what disappears is handled cleanly. It hurts when it multiplies empty pages to capture keywords, which Google classifies as spam.

In short

  • Generating pages in bulk is not a problem in itself: Google sanctions pages produced to manipulate rankings, with no value for the reader.
  • A programmatic page holds up when it carries real data the visitor will not find in that form elsewhere: a price, an area, an availability.
  • The real trap of a living catalogue is what leaves it: a vanished listing returns a 410 with useful alternatives, never a page served with a 200 or a redirect to the home page.
  • Check every language separately, templates, agreement and links included, or one error repeats across thousands of pages.

What programmatic SEO covers

Programmatic SEO is producing pages in bulk from a database and a template. One page per product, one per city, one per category and location combination, all in several languages. A catalogue of a few hundred items quickly becomes several thousand URLs.

There is nothing suspicious about the idea: it is how every classifieds, travel or ecommerce site works. What is suspicious is how some people use it: producing pages for keywords rather than for people. Google has a written rule on this. Its spam policy defines scaled content abuse as generating many pages for the primary purpose of manipulating search rankings and not helping users. Among its examples: using generative AI tools to produce many pages without adding value, stitching together content taken from other sites, including through automated translation, and creating pages that make little sense to a reader but contain search keywords.

So the line does not run between “generated” and “written by hand”. It runs between a page that serves useful data and a page that only exists to exist.

Our ground: a real estate catalogue in six languages

Palmora Property is our own real estate agency site in Phuket. Its catalogue of listings is synced every day from the agency’s CRM and published in six languages: English, French, Thai, Russian, Chinese and German. With the category pages and zone pages, also in six languages, the site declares more than 3,000 URLs.

None of that volume is written by hand. And it is by running it that we learned what separates a healthy programmatic catalogue from a page factory.

What works: real data on every page

A listing page answers a precise search with information the visitor will not find in that form anywhere else: price, area, zone, photos, availability. The page exists because the property exists. That is the healthiest form of programmatic SEO: the database is the page’s reason to exist, not a pretext. It is also what generative engines most readily pick up, precise and verifiable data, as we explain in the article on getting cited by ChatGPT.

Zone listing pages are trickier. Our audit of 22 July said it plainly: every page had a unique title and description, with the number of listings in the zone, but no editorial content, no subheading, and a main heading reduced to the zone name. Yet these are the pages that answer the most commercial queries. A grid of cards without a word of context is still a thin page.

So we wrote editorial content for eight zones (Rawai, Chalong, Bang Tao, Cherng Talay, Patong, Kata, Karon, Thalang): the market, the neighbourhoods, nearby zones, links to blog articles. It is written in English, then translated into the other languages with an AI model, with one rule: links to articles are never machine translated. They are rebuilt from the real articles of each language, and an article that does not exist in a language is simply dropped. A translated page that points to a missing article is exactly the content that “makes little sense to a reader”.

What also works: templates that are truly localized

A multilingual template is more than translated words. The classic trap: building the heading by gluing the property type to the contract, which gives “Villas Zu verkaufen” in German, and leaving the title and description tags in English in the languages nobody on the team reads. None of it shows when you proofread the French version.

The fix is to take those strings out of the code. On Palmora, each language has its own templates for the title, the description and the empty-list message, with the noun forms for each property type. Agreement with a number too: Russian has three forms depending on the count, and the results counter applies the same rule on first render and after a filter change. Same standard for listing titles, which are genuinely translated rather than copied in English into the other languages.

A language is not just its routes either. One forgotten string table is enough to serve an empty page with a 200 status that nobody notices. You have to search everywhere a list of languages is hard-coded, and test every language, page by page.

What hurts: pages that die badly

The real trap of a living catalogue is what leaves it. Every listing sold or unpublished disappears from the data, so its page disappears. But the URL is still indexed, shared, bookmarked. Left untreated, it falls into a raw 404, and the 404s pile up in Search Console as the catalogue turns over.

The usual answers do not fit:

  • A nice page served with a 200 saying “this property is no longer available”: to Google that is a soft 404, a new error in the console.
  • A 301 redirect per listing: there is no reliable target, the property no longer exists, and a redirect table you have to maintain always ends up pointing nowhere.

What we put in place: every URL of a vanished listing returns 410, meaning “gone for good”, with a genuinely useful page in the URL’s language. It suggests six similar listings, in tiers: same zone, same type and same contract within plus or minus 20% of the last known price, then widening. The memory of vanished listings is kept by the daily sync, and a listing put back online leaves it automatically. For listings marked “price on request”, which have no reference price, the page falls back to zone and type.

Two other traps of the same kind, worth checking on any catalogue:

  • Unknown URLs sent to the home page. Redirecting an invalid category to the home page turns every broken URL into one more redirect in the console. An unknown URL should return a 404.
  • The 404 page in the wrong language. After an internal rewrite to the error page, a site that reads the language from the error page’s URL, instead of the requested one, serves an English 404 in every language.

What also hurts: signals that lie

The sitemaps of a programmatic site are generated, so they can be wrong in bulk. Our July audit found three issues: a news sitemap that returned 404 when empty while being declared in robots.txt, a sitemap index whose last-modified date was the request time, and no date on the listing pages even though the CRM knows when each one was updated. An engine that notices a date lying learns to ignore it. All three are fixed: empty sitemap served with a 200, an index with no fabricated date, and the real update date on every listing.

Language links are the same kind of signal. An hreflang that announces a translation that does not exist, or that keeps the path of the starting language, sends the engine to a 404. We cover these cases in the article on hreflang mistakes.

One last rule applies to all generated content: nothing goes live without human review. A text produced by a model can contain an invented claim, and at catalogue scale, a template error repeats itself across thousands of pages.

The rule we apply

Before generating a family of pages, we ask four questions:

  1. Does every page answer a real search, with data the visitor will not find on the next page?
  2. What happens when the item disappears? A useful 410, a clean 404, never a redirect to the home page or an empty page served with a 200.
  3. Is the page right in every language, templates, agreement and links included, and not just translated?
  4. Are the signals true: sitemap dates, language links, canonicals?

If a single answer is no, it is better to generate fewer pages. That is what an SEO audit checks on a catalogue site, and what our SEO and GEO engine then monitors every week in production. Programmatic SEO is only one part of the job: we describe the rest in what an SEO agency should deliver in the age of AI, and the other articles on the subject are gathered on the SEO and AI visibility page.

Frequently asked questions

Does Google penalize programmatic SEO? Not as such. What Google sanctions is abuse: many pages produced for the primary purpose of manipulating rankings, without helping users. A catalogue where every page carries real, current and useful data is not concerned. The risk starts when pages all look alike, add nothing beyond the page next door, or only exist to target a keyword variant.

Should a removed listing return a 410 or a 301? A 301 if a page truly replaces it, for example the same product at another URL. Otherwise a 410, ideally with a page that suggests close alternatives, served with the right status code and never with a 200. Redirecting every removed listing to the home page or to a category amounts to manufacturing soft 404s, which Google ends up treating as errors.

Can generated pages be machine translated? Yes, if the translation is reviewed on its templates and grammatical agreement, and if links are rebuilt from the pages that actually exist in each language. Translating to multiply URLs, with no value for the reader, is explicitly listed by Google among the abuses. The right question is not who translated the page, but who the page serves.

Jim · founder of Brain Plus AI

Also in seo and ai visibility

  • Practical guide

    Hreflang and multilingual SEO: the costly mistakes

    Hreflang pointing to pages that do not exist, a poorly chosen x-default, 404s in the wrong language, untranslated titles, a language routed but forgotten elsewhere: the mistakes we found on our own multilingual sites, and how to avoid them.

    6min
  • Practical guide

    How to get cited by ChatGPT and AI search engines

    What actually gets a site cited by ChatGPT, Claude or Perplexity: let the bots read it, answer straight away, source every fact, mark up only what is shown. And how to measure it without fooling yourself.

    7min
  • Understand

    llms.txt: what it is really for

    The llms.txt file promises to help AI read your site. What the proposal says, what can be verified about real usage, what Google, OpenAI and Anthropic say about it, and why we publish one anyway.

    6min
WhatsApp