Skip to main content

I Had an AI Agent Audit My Self-Hosted Blog for Setup Leaks

How an AI coding agent audited my self-hosted WordPress blog for setup leaks in REST fields, rendered HTML and images, fixed them safely, and now gates new posts.

A few of my build notes said more than they should have. One named the hosting panel behind a side project, another mentioned the exact hardware in my homelab, and a couple of image filenames carried project names that the post body never did. None of it was a password, but together it was a map of my setup, and I don’t want that on a public blog.

🎯 Not sure if this will run on your hardware?Use our free Local LLM Hardware Checker — pick your GPU and RAM, see which models will run with real tokens/sec estimates.
Check my hardware →

So I gave the clean-up to an AI coding agent working from a sandboxed Linux box with a shell, curl and the same WordPress REST access I’d use myself. It audited every post in this section, wrote me a table of what leaked and where, and only changed things after I approved the list. This is how that worked and the checklist I now run before anything goes live.

The rule I gave the agent

Public posts can describe how something is built, never where it runs. That means no domain names for my other projects, no business or client names, no server or panel names, no IPs, no hosting provider, no hardware models and no plugin or file names that point back to a specific site. Code samples use example.com, and descriptions stay generic, like “a managed nginx panel” or “a small always-on box”.

A rule that loose needs a concrete term list, so I wrote one in a private file the agent reads and the blog never sees. It holds my project names, panel names, hardware models and so on. The scanner below only ever loads it from disk.

Advertisement

Why scanning the post body isn’t enough

The first pass looked at post content only, and it missed half the problems. A WordPress post leaks through a lot more than its body:

  • Title and slug. If the slug names a product or a panel, the URL itself leaks, and a body scrub doesn’t fix it.
  • Excerpt and meta description. If the excerpt is empty, many themes build the description from the first paragraph, so an old sentence can live on in og:description.
  • Featured image filename, title and alt text. my-project-panel-cover.png ends up in og:image and twitter:image on every share.
  • Tags. A tag named after a hardware model still has a public archive page, even after you remove it from every post.
  • Sitewide template parts. Footers and “more from this category” widgets show up on every page, so one leaky title in a widget appears under every clean post.
  • Response headers. Some panels and cache layers add their own header to every response, which tells anyone with curl -I what you run.

So the agent switched to scanning three layers for every post: the REST fields, the rendered HTML, and the media attached to it.

The scan the agent ran

This is the shape of it, trimmed down. The term list sits in a local file, one term per line.

#!/usr/bin/env bash
# leak-scan.sh: scan every published post in one category for private terms
SITE="https://example.com"
CAT=123                       # category ID
TERMS="$HOME/.private/leak-terms.txt"

curl -s "$SITE/wp-json/wp/v2/posts?categories=$CAT&per_page=100&_fields=id,link,featured_media" |
jq -r '.[] | "\(.id) \(.link) \(.featured_media)"' |
while read -r id link media; do
  {
    # 1. REST fields: title, slug, excerpt, content
    curl -s "$SITE/wp-json/wp/v2/posts/$id?_fields=title,slug,excerpt,content"
    # 2. Rendered page: meta description, og:*, twitter:*, widgets
    curl -s "$link"
    # 3. Featured media: filename, title, alt text
    curl -s "$SITE/wp-json/wp/v2/media/$media?_fields=source_url,title,alt_text"
  } | grep -oiFf "$TERMS" | sort | uniq -c | sed "s|^|$id |"
done

A hit in the rendered page but not in the REST fields means the leak is in a template part, not the post. The agent kept those in a separate column, because the fix lives somewhere else.

It also checked headers and old URLs, since those don’t show up in any HTML:

curl -sI https://example.com/ | grep -iE '^x-'      # panel or cache headers
curl -sI https://example.com/2026/09/old-slug/ | grep -iE '^(HTTP|location|x-redirect-by)'

What the agent did with the results

The audit came back as one table: post, URL, what leaked, whether the post had a real AI angle, and the fix it proposed. I made a second rule at the same time, that every post here has to combine AI with self-hosting or dev tools, so the table checked that too. I approved the fixes in batches, and the agent worked through the same loop on each one:

  1. Back up first. It saved each post’s full edit view (?context=edit) and its raw HTML to a local folder before writing anything, so rolling back is a single REST update.
  2. Light scrub or full rewrite. If a post had one stray hardware mention, it just swapped in a generic phrase. If the leak was in the title or slug, or the whole post was built around a named project, it rewrote the post with a new title, slug and excerpt.
  3. New slug, keep the old link working. WordPress remembers old slugs and 301s them to the new URL on its own, which the agent confirmed with the x-redirect-by: WordPress header. One post was unpublished rather than rewritten, and that old URL needed a proper redirect rule from a redirect plugin.
  4. Replace leaky images, don’t rename them. It uploaded a fresh cover under a neutral, date-based filename, switched the featured image, checked the rendered og:image, and only then deleted the old media item.
  5. Clean up the leftovers. It deleted tags named after hardware once they were empty, and updated internal links so no post pointed at an old slug.
  6. Re-scan. Every changed post had to return 200, every old slug had to return 301, and the term scan had to come back empty for the post itself.

What surprised me

  • Deleted isn’t gone. When you delete an image, the server returns 404, but a CDN in front can keep serving its cached copy for weeks. The agent checked the CDN’s cache-status header on each old URL and handed me a list of URLs to purge instead of assuming they were gone.
  • SEO meta can be silently ignored. The publish script was writing SEO-plugin meta keys for a plugin that wasn’t installed on that site. WordPress accepted the request and did nothing with them. The fix was to set the title, excerpt and featured image the theme actually uses, then verify the rendered tags rather than trusting the API response.
  • Not every leak sits in the database. One block was printed by PHP in the footer. The agent ruled out widgets, menus and template parts over REST, then briefly switched plugins off one at a time and fetched an uncached copy of the page to see whether the block disappeared. When none of them was the source, it stopped and told me the fix needed file access, instead of guessing.
  • GUIDs keep old slugs. WordPress never changes a post’s GUID. It’s only visible in the API, so I left it, but it’s worth knowing that a rename isn’t a total erase.

The pre-publish check I run now

The daily publishing job for this section runs the same scan before it publishes anything. It checks the rendered title, body, meta description, OG and Twitter tags, alt text and image filenames against the private term list, and the post doesn’t go out until the scan comes back clean. It’s the same idea as my 7 rules for letting an AI agent maintain your homelab: the agent can move fast, but it has to prove each change with a request I can re-run.

Ship checklist

  1. Write your private term list down (projects, domains, panels, hosts, hardware, client names) and keep it off the blog.
  2. Scan REST fields, rendered HTML and media for every post, not just the content.
  3. Back up each post’s edit view before an agent changes anything.
  4. If the title or slug leaks, rewrite the post with a new slug and confirm the old one 301s.
  5. Upload replacement images under neutral filenames, switch the featured image, then delete the old one.
  6. Delete empty tags that name hardware or providers, and fix internal links that point at old slugs.
  7. Check response headers and sitewide template parts, because they leak on every page.
  8. Purge the CDN for deleted image URLs.
  9. Run the same scan automatically before every new post.
Advertisement

FAQ

Isn’t this overkill for a personal blog?

Each detail is harmless on its own. Put them together, though (panel, host, hardware, which sites share a server), and you’ve given anyone a head start on recon. I work in security, so I’d rather publish the method and keep the map private.

Did the agent change posts on its own?

No. The audit was read-only. The agent proposed a fix for each post, I approved them in batches, and it backed up every post before writing. Deleting media, tags or posts waited for my yes.

Can a local model do this instead of a cloud agent?

The scan itself is just curl, jq and grep, so any model that can run shell commands can drive it. The model is mostly useful for the rewrites. If you’d rather keep your term list and drafts on your own hardware, my local LLM homelab stack post covers what I keep running.

Closest related Builders notes?

I Let an AI Agent Ship Live WordPress Pages and a www Redirect uses the same back-up, dry-run and verify loop on a live site, and How I Run Claude Code 24/7 on an Always-On Home Box With tmux covers where an agent like this lives between jobs.

Written by MustafaEngineer running a 24/7 homelab since 2022. Every guide here is built and tested on my own hardware. No paid placements.

Keep reading