Skip to main content

My AI Agent Read the Terms Before Building a 24/7 Monitor

How an AI agent split feasible from authorised before automating my self-hosted listing app, picked the site's own email alerts, and set a watch to prove them first.

I have a small self-hosted app at home that scores second-hand listings for the few things I actually shop for. Today it only sees a listing when I press a capture button in my own browser. This week I asked an AI agent to turn it into a 24/7 monitor: new listings, price drops, alerts, the lot.

The agent didn’t start with code. It started with the marketplace’s terms of use, and that changed the whole build.

The ask, and the rule I set first

The wish list was a proper event-driven adapter: daily discovery of my saved searches, new-listing and price-change events, idempotent processing, retries, health checks and price history. I also asked it to evaluate changedetection.io, the self-hosted page watcher, as the engine.

Before any of that, I gave it one hard rule: no bypassing CAPTCHAs, bot challenges, logins or any other access control. And keep “technically possible” separate from “actually allowed” in every answer.

Advertisement

Research from primary sources only

The agent split the job in two and wrote both into one research brief:

  • The tool. It cloned changedetection.io and read the source, the OpenAPI spec and the wiki instead of blog posts. That covered fetchers, per-watch filters, webhooks and what needs a headless browser container.
  • The site. It kept contact with the marketplace to the bare minimum: one request each for robots.txt and the homepage, and one attempt at the terms page. It didn’t fetch a single search or listing page.

Those few requests were enough. The homepage answered with a JavaScript bot-challenge page, not content. The terms, once read, explicitly ban bots, scripts and crawlers that scrape the platform or build a database from it.

Feasible is not the same as authorised

This is the part I liked most. The brief has a table with two separate columns for every option: does it work technically, and is it authorised. The page watcher pointed at search results was a clear yes in the first column. It could watch a results URL once a day, trim the page down to listing cards, ignore ads and “2 hours ago” timestamps, and POST a JSON webhook to my app.

It was also a clear no in the second. Running it would break the terms, and in practice it would mean defeating the bot challenge, which my rule forbids. So the agent wrote down what permission would actually look like instead of building it anyway: written approval from the marketplace that names the monitoring, its scope and rate, and the agent and IP, plus a data licence. Ideally that comes with an allowlisted client or an official feed.

The authorised path: the site’s own alerts

The clean option was boring, which is usually a good sign. The marketplace already offers saved-search alerts by email. Those land in my own inbox, and reading my own inbox is fine. The plan became:

  1. I create saved searches with email alerts in the site’s own UI.
  2. A read-only inbox reader picks up only genuine alert emails.
  3. Each alert becomes a signed webhook into my self-hosted app, which values and scores it.
  4. Per-listing detail still comes through my manual capture button.

Nothing in that pipeline makes an automated request to the marketplace.

saved-search alert email  ->  read-only inbox check
  ->  is it a real alert? (not marketing)
  ->  signed webhook  ->  self-hosted app: dedupe, value, score, notify

Old emails said “don’t trust it yet”

Before building the reader, the agent looked for evidence in my mailbox, read-only. It saved sanitised samples of every relevant type it found and compared them.

  • The older saved-search digests weren’t consistent. Some had direct listing links with an ID and a price. Others hid everything behind click-tracking redirects.
  • The digests said “past 24 hours” but didn’t arrive every day.
  • The newest emails in the inbox were marketing reminders (“don’t miss out on these deals”), not alerts, and their links went only through a tracker.

That was enough for my second rule. An email integration doesn’t count as working until a real alert has arrived and been checked. Marketing mail that looks similar doesn’t count.

A watch with a deadline instead of a build

So I created three saved searches, and the agent set up a small scheduled check instead of a monitor. Every hour during the day, it looks at my inbox for the first real alert from each search. It records whether that alert has a listing ID, a price and a direct link my app can use, and it logs and ignores the marketing reminders.

It also has a deadline. If any search is still silent after a week, the watch reports which ones, and email alerts get marked as not a viable source. As I write this, it has logged a day of checks and no real alert yet. My existing app version stays exactly as it is, and the always-on box it runs on hasn’t been touched.

That’s the opposite of how I’d normally build. It would have been easy to ship a polished adapter with webhooks, retries and a dashboard that never receives a single event. I’d rather wait for real data first.

Ship checklist

  1. Write the access rule down before research starts: no bypassing challenges, logins or controls.
  2. Have the agent read the target site’s terms and robots.txt with as few requests as possible, and no listing pages.
  3. Read the self-hosted tool’s source and docs, not summaries of them.
  4. Give every option two columns: technically feasible, and authorised.
  5. Write down what permission would actually be needed for the blocked options.
  6. Prefer the site’s own notification features, received in an inbox you own.
  7. Sample old emails read-only and redact recipients and tracking tokens.
  8. Treat an integration as unproven until a real event arrives and parses.
  9. Run a scheduled watch with a deadline before building the adapter.
  10. Leave the working version and its host alone until there’s a viable path.

FAQ

Isn’t one robots.txt request already a bit grey?

Fetching robots.txt and a terms page is what they’re there for. The point was to read the rules with as little traffic as possible. The agent made no requests for search results or listings, and it stopped as soon as it hit the challenge page.

Why not run the page watcher with a headless browser? It would probably get through.

“Probably gets through” is exactly the bot-challenge bypass I ruled out, and the terms forbid it anyway. Technically possible doesn’t make it allowed. If the marketplace ever grants written access or offers a feed, the page watcher setup from the brief is ready to use.

What if the alert emails never show up?

Then the honest answer is that there’s no suitable automatic source, and the app keeps using the manual capture button. The watch exists so that answer comes from a week of real observations, not a guess.

Where does the AI actually help here?

It did the reading I’d skip: tool source code, an API spec, a long terms page and years of old email, all checked against the rules I set. Then it turned that into a decision table and a small scheduled check. The code it didn’t write mattered as much as the code it did.

Closest related Builders notes?

My AI Agent Found an SSH Key and Chose Not to Use It is the same instinct applied to credentials. 7 Rules for Letting an AI Agent Maintain Your Homelab covers the guardrails, and I Gave My AI Agents One Shared Changelog With a Rollback Line is how changes like this get logged.

Written by MustafaEngineer running a 24/7 homelab since 2022. Every guide here is built and tested on my own hardware. No paid placements.

Keep reading