Add implementation plan for Playwright fallback
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013gwc3aCiyQxmCWQke8RvEz
This commit is contained in:
@@ -0,0 +1,633 @@
|
||||
# Korting Playwright-fallback Implementation Plan
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||
|
||||
**Goal:** When the existing plain-fetch price scraper fails (network error or no price found), retry once with a real headless-Chromium page load (Playwright) before recording the check as failed.
|
||||
|
||||
**Architecture:** A new, self-contained `src/browserScraper.js` module wraps Playwright and reuses `extractPrice()` from the existing `src/scraper.js` (no duplicated parsing logic). `checkProducts.js` gains one new *optional* injected dependency that it calls only when the primary fetch already failed — `scraper.js` itself is untouched. `server.js` lazily launches one browser per `runCheck()` run (only if a fallback is actually needed), reuses it for the rest of that run, and always closes it in a `finally`.
|
||||
|
||||
**Tech Stack:** Playwright (`chromium`), same Node.js 22 / CommonJS / `node --test` stack as the rest of the project.
|
||||
|
||||
**Spec:** `docs/superpowers/specs/2026-09-05-korting-playwright-fallback-design.md` (builds on `docs/superpowers/specs/2026-09-05-korting-prijstracker-design.md`)
|
||||
|
||||
## Global Constraints
|
||||
|
||||
- Node.js 22, CommonJS modules, no TypeScript — matches the existing codebase.
|
||||
- `src/scraper.js` and its tests are **not modified** by this plan — the fallback is a separate module, invoked from `checkProducts.js`, not from inside `fetchAndExtractPrice`.
|
||||
- The browser fallback is invoked **only** when the primary `fetchAndExtractPrice` result has `error` or `price === null` — never when it already found a price.
|
||||
- Every existing test must keep passing unmodified in behavior (the new dependency is optional; omitting it from `deps` must not change any existing test's outcome).
|
||||
- `playwright` is a production dependency (not `devDependencies`) — the Docker runtime stage needs it, not just the build stage.
|
||||
- **Local `npm install` must not attempt to download the actual Chromium binary** — set `PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1` in the environment before running `npm install` so the postinstall step is skipped. The real browser binary is only installed inside the Docker image (Task 4) via `npx playwright install --with-deps chromium`. Tasks 1-3 never need a real Chromium binary in the dev/test sandbox — all automated tests use fake browser/page objects (dependency injection), and Task 3's manual check deliberately avoids exercising the fallback path for this reason. Task 4's Docker build/run verification is the first (and only) point in this plan where a real browser actually launches.
|
||||
- Timeout for a fallback page load: 20000ms (`timeoutMs` default), passed to Playwright's `page.goto` timeout option.
|
||||
|
||||
---
|
||||
|
||||
### Task 1: Browser-based scraper module
|
||||
|
||||
**Files:**
|
||||
- Modify: `package.json` (add `playwright` dependency)
|
||||
- Create: `src/browserScraper.js`
|
||||
- Create: `test/browserScraper.test.js`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `extractPrice` from `src/scraper.js` (already exported, unchanged signature: `extractPrice(html) -> { price: number|null, method: string|null }`).
|
||||
- Produces:
|
||||
- `launchBrowser() -> Promise<Browser>` — thin wrapper around Playwright's `chromium.launch({ headless: true })`.
|
||||
- `fetchAndExtractPriceViaBrowser(browser, url, { timeoutMs = 20000 } = {}) -> Promise<{ price: number|null, method: string|null, error: Error|null }>` — same return shape as `scraper.js`'s `fetchAndExtractPrice`, so later tasks can treat both interchangeably.
|
||||
|
||||
- [ ] **Step 1: Add the `playwright` dependency**
|
||||
|
||||
Run (this environment must not attempt to download the Chromium binary — see Global Constraints):
|
||||
|
||||
```bash
|
||||
PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1 npm install playwright
|
||||
```
|
||||
|
||||
Expected: `package.json`'s `dependencies` gains a `playwright` entry (not `devDependencies`), `package-lock.json` updates, and the command completes without trying to download a browser binary.
|
||||
|
||||
- [ ] **Step 2: Write the failing tests**
|
||||
|
||||
Create `test/browserScraper.test.js`:
|
||||
|
||||
```js
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { fetchAndExtractPriceViaBrowser } = require('../src/browserScraper');
|
||||
|
||||
test('fetchAndExtractPriceViaBrowser extracts a price from the rendered page', async () => {
|
||||
let closed = false;
|
||||
const html = '<html><head><meta property="og:price:amount" content="42.00" /></head><body></body></html>';
|
||||
const page = {
|
||||
goto: async () => {},
|
||||
content: async () => html,
|
||||
close: async () => { closed = true; },
|
||||
};
|
||||
const browser = { newPage: async () => page };
|
||||
const result = await fetchAndExtractPriceViaBrowser(browser, 'https://example.com');
|
||||
assert.deepEqual(result, { price: 42, method: 'meta', error: null });
|
||||
assert.equal(closed, true);
|
||||
});
|
||||
|
||||
test('fetchAndExtractPriceViaBrowser returns an error when navigation fails, and still closes the page', async () => {
|
||||
let closed = false;
|
||||
const page = {
|
||||
goto: async () => { throw new Error('Timeout 20000ms exceeded'); },
|
||||
content: async () => '',
|
||||
close: async () => { closed = true; },
|
||||
};
|
||||
const browser = { newPage: async () => page };
|
||||
const result = await fetchAndExtractPriceViaBrowser(browser, 'https://example.com');
|
||||
assert.equal(result.price, null);
|
||||
assert.equal(result.method, null);
|
||||
assert.ok(result.error instanceof Error);
|
||||
assert.equal(closed, true);
|
||||
});
|
||||
|
||||
test('fetchAndExtractPriceViaBrowser returns null price when nothing matches, and still closes the page', async () => {
|
||||
let closed = false;
|
||||
const page = {
|
||||
goto: async () => {},
|
||||
content: async () => '<html><body><p>Geen prijs</p></body></html>',
|
||||
close: async () => { closed = true; },
|
||||
};
|
||||
const browser = { newPage: async () => page };
|
||||
const result = await fetchAndExtractPriceViaBrowser(browser, 'https://example.com');
|
||||
assert.deepEqual(result, { price: null, method: null, error: null });
|
||||
assert.equal(closed, true);
|
||||
});
|
||||
|
||||
test('fetchAndExtractPriceViaBrowser closes the page even if reading its content throws', async () => {
|
||||
let closed = false;
|
||||
const page = {
|
||||
goto: async () => {},
|
||||
content: async () => { throw new Error('content failed'); },
|
||||
close: async () => { closed = true; },
|
||||
};
|
||||
const browser = { newPage: async () => page };
|
||||
const result = await fetchAndExtractPriceViaBrowser(browser, 'https://example.com');
|
||||
assert.ok(result.error instanceof Error);
|
||||
assert.equal(closed, true);
|
||||
});
|
||||
|
||||
test('fetchAndExtractPriceViaBrowser passes the timeout option through to page.goto', async () => {
|
||||
let receivedOptions;
|
||||
const page = {
|
||||
goto: async (url, options) => { receivedOptions = options; },
|
||||
content: async () => '<html><body></body></html>',
|
||||
close: async () => {},
|
||||
};
|
||||
const browser = { newPage: async () => page };
|
||||
await fetchAndExtractPriceViaBrowser(browser, 'https://example.com', { timeoutMs: 5000 });
|
||||
assert.equal(receivedOptions.timeout, 5000);
|
||||
});
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Run tests to verify they fail**
|
||||
|
||||
Run: `npm test`
|
||||
Expected: FAIL with "Cannot find module '../src/browserScraper'"
|
||||
|
||||
- [ ] **Step 4: Write `src/browserScraper.js`**
|
||||
|
||||
```js
|
||||
const { chromium } = require('playwright');
|
||||
const { extractPrice } = require('./scraper');
|
||||
|
||||
async function launchBrowser() {
|
||||
return chromium.launch({ headless: true });
|
||||
}
|
||||
|
||||
async function fetchAndExtractPriceViaBrowser(browser, url, { timeoutMs = 20000 } = {}) {
|
||||
let page;
|
||||
try {
|
||||
page = await browser.newPage();
|
||||
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
|
||||
const html = await page.content();
|
||||
const { price, method } = extractPrice(html);
|
||||
return { price, method, error: null };
|
||||
} catch (err) {
|
||||
return { price: null, method: null, error: err };
|
||||
} finally {
|
||||
if (page) {
|
||||
await page.close().catch(() => {});
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = { launchBrowser, fetchAndExtractPriceViaBrowser };
|
||||
```
|
||||
|
||||
- [ ] **Step 5: Run tests to verify they pass**
|
||||
|
||||
Run: `npm test`
|
||||
Expected: PASS (all `browserScraper.test.js` tests + every previously-passing test, unchanged)
|
||||
|
||||
- [ ] **Step 6: Commit**
|
||||
|
||||
```bash
|
||||
git add package.json package-lock.json src/browserScraper.js test/browserScraper.test.js
|
||||
git commit -m "feat: add Playwright-based fallback scraper"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 2: Wire the fallback into the price-check orchestration
|
||||
|
||||
**Files:**
|
||||
- Modify: `src/checkProducts.js`
|
||||
- Modify: `test/checkProducts.test.js`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `fetchAndExtractPriceViaBrowser`-shaped function (Task 1's return shape), passed as the new optional `deps.fetchAndExtractPriceViaBrowser(url) -> Promise<{price, method, error}>`. `checkProducts.js` does not `require` Task 1's module directly — the composition root (Task 3) curries in the real browser instance.
|
||||
- Produces: `checkProduct`/`checkAllProducts` — same exported names and signatures as before; the only change is that `checkProduct`'s deps object now recognizes one more optional key.
|
||||
|
||||
The current `src/checkProducts.js` (for reference — this is what you are editing):
|
||||
|
||||
```js
|
||||
async function checkProduct(product, deps) {
|
||||
const { fetchAndExtractPrice, sendMail, recordCheckSuccess, recordCheckFailure, now } = deps;
|
||||
const result = await fetchAndExtractPrice(product.url);
|
||||
|
||||
if (result.error) {
|
||||
console.error(`[check] fetch_error "${product.name}" (${product.url}): ${result.error.message}`);
|
||||
recordCheckFailure(product.id, 'fetch_error');
|
||||
return { status: 'fetch_error' };
|
||||
}
|
||||
if (result.price == null) {
|
||||
console.error(`[check] not_found "${product.name}" (${product.url}): geen prijs gevonden op de pagina`);
|
||||
recordCheckFailure(product.id, 'not_found');
|
||||
return { status: 'not_found' };
|
||||
}
|
||||
|
||||
const isBaseline = product.last_price == null;
|
||||
const changed = !isBaseline && result.price !== product.last_price;
|
||||
|
||||
if (changed) {
|
||||
try {
|
||||
await sendMail({
|
||||
name: product.name,
|
||||
url: product.url,
|
||||
oldPrice: product.last_price,
|
||||
newPrice: result.price,
|
||||
});
|
||||
} catch (err) {
|
||||
console.error(`[check] mail_error "${product.name}" (${product.url}): ${err && err.message}`);
|
||||
recordCheckFailure(product.id, 'mail_error');
|
||||
return { status: 'mail_error' };
|
||||
}
|
||||
}
|
||||
|
||||
recordCheckSuccess(product.id, { price: result.price, checkedAt: now() });
|
||||
return { status: 'ok', changed };
|
||||
}
|
||||
|
||||
async function checkAllProducts(products, deps) {
|
||||
const results = [];
|
||||
for (const product of products) {
|
||||
try {
|
||||
const result = await checkProduct(product, deps);
|
||||
results.push({ productId: product.id, ...result });
|
||||
} catch (err) {
|
||||
console.error(`[check] error "${product.name}" (${product.url}): ${err && err.message}`);
|
||||
results.push({ productId: product.id, status: 'error', error: err });
|
||||
}
|
||||
}
|
||||
return results;
|
||||
}
|
||||
|
||||
module.exports = { checkProduct, checkAllProducts };
|
||||
```
|
||||
|
||||
The current `test/checkProducts.test.js` (for reference — you are appending new tests to this file, not replacing the existing ones):
|
||||
|
||||
```js
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { checkProduct, checkAllProducts } = require('../src/checkProducts');
|
||||
|
||||
function makeDeps(overrides = {}) {
|
||||
const failures = [];
|
||||
const successes = [];
|
||||
const mailCalls = [];
|
||||
const deps = {
|
||||
fetchAndExtractPrice: async () => ({ price: 10, method: 'meta', error: null }),
|
||||
sendMail: async (payload) => { mailCalls.push(payload); },
|
||||
recordCheckFailure: (id, status) => failures.push({ id, status }),
|
||||
recordCheckSuccess: (id, data) => successes.push({ id, ...data }),
|
||||
now: () => '2026-09-05T08:00:00.000Z',
|
||||
...overrides,
|
||||
};
|
||||
return { deps, failures, successes, mailCalls };
|
||||
}
|
||||
|
||||
// Failure paths log to stderr; capture it so the test output stays clean
|
||||
// while still proving the logging happened.
|
||||
function captureErrorLog(t) {
|
||||
const logged = [];
|
||||
t.mock.method(console, 'error', (...args) => { logged.push(args.join(' ')); });
|
||||
return logged;
|
||||
}
|
||||
|
||||
// ... (existing tests, unchanged — do not modify or remove any of them)
|
||||
```
|
||||
|
||||
- [ ] **Step 1: Write the failing tests**
|
||||
|
||||
Append these tests to the end of `test/checkProducts.test.js` (after the existing `checkAllProducts processes every product independently...` test, before the final closing of the file):
|
||||
|
||||
```js
|
||||
test('checkProduct tries the browser fallback when the primary fetch errors, and uses its result', async (t) => {
|
||||
captureErrorLog(t);
|
||||
const browserCalls = [];
|
||||
const { deps, successes, mailCalls } = makeDeps({
|
||||
fetchAndExtractPrice: async () => ({ price: null, method: null, error: new Error('down') }),
|
||||
fetchAndExtractPriceViaBrowser: async (url) => {
|
||||
browserCalls.push(url);
|
||||
return { price: 10, method: 'json-ld', error: null };
|
||||
},
|
||||
});
|
||||
const product = { id: 7, name: 'X', url: 'https://x', last_price: null };
|
||||
const result = await checkProduct(product, deps);
|
||||
assert.equal(result.status, 'ok');
|
||||
assert.deepEqual(browserCalls, ['https://x']);
|
||||
assert.equal(successes.length, 1);
|
||||
assert.equal(mailCalls.length, 0);
|
||||
});
|
||||
|
||||
test('checkProduct tries the browser fallback when the primary fetch finds no price', async (t) => {
|
||||
captureErrorLog(t);
|
||||
const browserCalls = [];
|
||||
const { deps, successes } = makeDeps({
|
||||
fetchAndExtractPrice: async () => ({ price: null, method: null, error: null }),
|
||||
fetchAndExtractPriceViaBrowser: async (url) => {
|
||||
browserCalls.push(url);
|
||||
return { price: 12, method: 'regex', error: null };
|
||||
},
|
||||
});
|
||||
const product = { id: 8, name: 'X', url: 'https://x', last_price: null };
|
||||
const result = await checkProduct(product, deps);
|
||||
assert.equal(result.status, 'ok');
|
||||
assert.deepEqual(browserCalls, ['https://x']);
|
||||
assert.equal(successes.length, 1);
|
||||
});
|
||||
|
||||
test('checkProduct never calls the browser fallback when the primary fetch already succeeds', async () => {
|
||||
const browserCalls = [];
|
||||
const { deps } = makeDeps({
|
||||
fetchAndExtractPriceViaBrowser: async (url) => {
|
||||
browserCalls.push(url);
|
||||
return { price: 999, method: 'meta', error: null };
|
||||
},
|
||||
});
|
||||
const product = { id: 9, name: 'X', url: 'https://x', last_price: null };
|
||||
await checkProduct(product, deps);
|
||||
assert.equal(browserCalls.length, 0);
|
||||
});
|
||||
|
||||
test('checkProduct records fetch_error when both the primary fetch and the browser fallback fail', async (t) => {
|
||||
const logged = captureErrorLog(t);
|
||||
const { deps, failures, successes } = makeDeps({
|
||||
fetchAndExtractPrice: async () => ({ price: null, method: null, error: new Error('down') }),
|
||||
fetchAndExtractPriceViaBrowser: async () => ({ price: null, method: null, error: new Error('still down') }),
|
||||
});
|
||||
const product = { id: 10, name: 'X', url: 'https://x', last_price: 10 };
|
||||
const result = await checkProduct(product, deps);
|
||||
assert.equal(result.status, 'fetch_error');
|
||||
assert.deepEqual(failures, [{ id: 10, status: 'fetch_error' }]);
|
||||
assert.equal(successes.length, 0);
|
||||
assert.equal(logged.length, 1);
|
||||
assert.match(logged[0], /fetch_error.*still down/);
|
||||
});
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Run tests to verify they fail**
|
||||
|
||||
Run: `npm test`
|
||||
Expected: FAIL — the 4 new tests fail because `checkProduct` doesn't yet call `fetchAndExtractPriceViaBrowser` (e.g. the "tries the browser fallback" tests fail their `browserCalls` assertions with an empty array; the "never calls" test passes vacuously already, which is fine — it's a guard for the next step, not expected to fail).
|
||||
|
||||
- [ ] **Step 3: Modify `src/checkProducts.js`**
|
||||
|
||||
Change only the first few lines of `checkProduct` (everything from the baseline/changed logic onward stays exactly as it is):
|
||||
|
||||
```js
|
||||
async function checkProduct(product, deps) {
|
||||
const {
|
||||
fetchAndExtractPrice, fetchAndExtractPriceViaBrowser,
|
||||
sendMail, recordCheckSuccess, recordCheckFailure, now,
|
||||
} = deps;
|
||||
let result = await fetchAndExtractPrice(product.url);
|
||||
|
||||
if ((result.error || result.price == null) && fetchAndExtractPriceViaBrowser) {
|
||||
result = await fetchAndExtractPriceViaBrowser(product.url);
|
||||
}
|
||||
|
||||
if (result.error) {
|
||||
console.error(`[check] fetch_error "${product.name}" (${product.url}): ${result.error.message}`);
|
||||
recordCheckFailure(product.id, 'fetch_error');
|
||||
return { status: 'fetch_error' };
|
||||
}
|
||||
if (result.price == null) {
|
||||
console.error(`[check] not_found "${product.name}" (${product.url}): geen prijs gevonden op de pagina`);
|
||||
recordCheckFailure(product.id, 'not_found');
|
||||
return { status: 'not_found' };
|
||||
}
|
||||
|
||||
const isBaseline = product.last_price == null;
|
||||
const changed = !isBaseline && result.price !== product.last_price;
|
||||
|
||||
if (changed) {
|
||||
try {
|
||||
await sendMail({
|
||||
name: product.name,
|
||||
url: product.url,
|
||||
oldPrice: product.last_price,
|
||||
newPrice: result.price,
|
||||
});
|
||||
} catch (err) {
|
||||
console.error(`[check] mail_error "${product.name}" (${product.url}): ${err && err.message}`);
|
||||
recordCheckFailure(product.id, 'mail_error');
|
||||
return { status: 'mail_error' };
|
||||
}
|
||||
}
|
||||
|
||||
recordCheckSuccess(product.id, { price: result.price, checkedAt: now() });
|
||||
return { status: 'ok', changed };
|
||||
}
|
||||
```
|
||||
|
||||
(`checkAllProducts` and `module.exports` are unchanged — leave them exactly as they are.)
|
||||
|
||||
- [ ] **Step 4: Run tests to verify they pass**
|
||||
|
||||
Run: `npm test`
|
||||
Expected: PASS (all new tests + every previously-passing test, unchanged)
|
||||
|
||||
- [ ] **Step 5: Commit**
|
||||
|
||||
```bash
|
||||
git add src/checkProducts.js test/checkProducts.test.js
|
||||
git commit -m "feat: fall back to the browser scraper when the primary fetch fails"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 3: Wire the browser lifecycle into the composition root
|
||||
|
||||
**Files:**
|
||||
- Modify: `server.js`
|
||||
|
||||
**Interfaces:**
|
||||
- Consumes: `launchBrowser`, `fetchAndExtractPriceViaBrowser` (Task 1); `checkProduct`'s new optional `fetchAndExtractPriceViaBrowser` deps key (Task 2).
|
||||
- Produces: nothing new for other tasks to consume — `server.js` is the top of the dependency graph.
|
||||
|
||||
No new automated test — this is a composition-root change, verified with a manual smoke check (same pattern as the original server.js task). **This step does not require a real Chromium binary**: the smoke check below only exercises a product whose plain fetch succeeds, so the lazy browser launch never actually triggers.
|
||||
|
||||
- [ ] **Step 1: Modify `server.js`**
|
||||
|
||||
Add the import (alongside the other `require`s near the top):
|
||||
|
||||
```js
|
||||
const { launchBrowser, fetchAndExtractPriceViaBrowser } = require('./src/browserScraper');
|
||||
```
|
||||
|
||||
Replace the `runCheck` function with a version that lazily launches one browser per run, only if a fallback is actually needed, and always closes it:
|
||||
|
||||
```js
|
||||
async function runCheck() {
|
||||
const products = listProducts(db);
|
||||
let browserPromise = null;
|
||||
function getBrowser() {
|
||||
if (!browserPromise) {
|
||||
browserPromise = launchBrowser();
|
||||
}
|
||||
return browserPromise;
|
||||
}
|
||||
try {
|
||||
const deps = {
|
||||
fetchAndExtractPrice,
|
||||
fetchAndExtractPriceViaBrowser: async (url) => {
|
||||
const browser = await getBrowser();
|
||||
return fetchAndExtractPriceViaBrowser(browser, url);
|
||||
},
|
||||
sendMail: (payload) => sendPriceChangeEmail(transport, {
|
||||
from: process.env.SMTP_FROM,
|
||||
to: process.env.NOTIFY_EMAIL,
|
||||
...payload,
|
||||
}),
|
||||
recordCheckFailure: (id, status) => recordCheckFailure(db, id, status),
|
||||
recordCheckSuccess: (id, data) => recordCheckSuccess(db, id, data),
|
||||
now: () => new Date().toISOString(),
|
||||
};
|
||||
const results = await checkAllProducts(products, deps);
|
||||
console.log(`[check] ${products.length} product(en) gecontroleerd: ${summarize(results)}`);
|
||||
return results;
|
||||
} finally {
|
||||
if (browserPromise) {
|
||||
const browser = await browserPromise;
|
||||
await browser.close().catch((err) => {
|
||||
console.error(`[check] kon de browser niet netjes sluiten: ${err && err.message}`);
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
(Everything else in `server.js` — the required-env check, `openDb`, `createTransport`, `summarize`, `startScheduler(runCheck)`, `createApp({...})`, `app.listen(...)` — stays exactly as it is.)
|
||||
|
||||
- [ ] **Step 2: Manual smoke check (no real browser needed for this check)**
|
||||
|
||||
Deliberately **do not** call `/check-now` in this step: with the Task 2
|
||||
change, any URL whose plain fetch doesn't yield a price (which is most
|
||||
URLs without a real product price, e.g. `https://example.com`) now
|
||||
triggers the browser fallback — and no real Chromium binary is installed
|
||||
in this dev sandbox (see Global Constraints). Exercising the actual
|
||||
fallback path with a real browser is Task 4's job, once Chromium is
|
||||
installed in the Docker image. This step only proves `server.js` still
|
||||
starts and its non-check routes still work after the edit — i.e. that the
|
||||
new `require('./src/browserScraper')` and the rewritten `runCheck` are
|
||||
syntactically sound and don't break anything at startup.
|
||||
|
||||
Run:
|
||||
```bash
|
||||
cp .env.example .env
|
||||
# edit .env: fill in real BASIC_AUTH_USER/PASS and SMTP_FROM/NOTIFY_EMAIL at minimum
|
||||
set -a; source .env; set +a
|
||||
node server.js
|
||||
```
|
||||
In another terminal:
|
||||
```bash
|
||||
curl -u "$BASIC_AUTH_USER:$BASIC_AUTH_PASS" -X POST http://localhost:3000/products \
|
||||
-d "name=Testproduct" -d "url=https://example.com"
|
||||
curl -u "$BASIC_AUTH_USER:$BASIC_AUTH_PASS" http://localhost:3000/
|
||||
```
|
||||
Expected: both succeed (302, then 200 with "Testproduct" visible in the
|
||||
list). Stop the server with Ctrl-C — do not proceed to `/check-now` here.
|
||||
|
||||
- [ ] **Step 3: Commit**
|
||||
|
||||
```bash
|
||||
git add server.js
|
||||
git commit -m "feat: launch a lazy per-run browser for the fallback scraper"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Task 4: Docker image — install Chromium for the fallback
|
||||
|
||||
**Files:**
|
||||
- Modify: `Dockerfile`
|
||||
- Modify: `README.md`
|
||||
|
||||
**Interfaces:** none — this task only changes the deployment image and documentation.
|
||||
|
||||
The current `Dockerfile` (for reference):
|
||||
|
||||
```dockerfile
|
||||
FROM node:22-slim AS build
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends python3 build-essential \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
WORKDIR /app
|
||||
COPY package.json package-lock.json ./
|
||||
RUN npm ci --omit=dev
|
||||
COPY . .
|
||||
|
||||
FROM node:22-slim
|
||||
WORKDIR /app
|
||||
COPY --from=build /app .
|
||||
ENV NODE_ENV=production
|
||||
# De SQLite-database hoort op het persistente volume te staan; zonder deze
|
||||
# default zou een vergeten DB_PATH stilletjes naar de container-fs schrijven
|
||||
# en bij elke redeploy alle data verliezen.
|
||||
ENV DB_PATH=/data/korting.db
|
||||
VOLUME /data
|
||||
EXPOSE 3000
|
||||
CMD ["node", "server.js"]
|
||||
```
|
||||
|
||||
- [ ] **Step 1: Modify the `Dockerfile`'s runtime stage**
|
||||
|
||||
Add the Chromium install line right after `COPY --from=build /app .` (this is the first point where `playwright` from `node_modules` is available to `npx`, and it's the stage that actually runs the browser, so the browser binary + its system libraries belong here, not in the build stage):
|
||||
|
||||
```dockerfile
|
||||
FROM node:22-slim AS build
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends python3 build-essential \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
WORKDIR /app
|
||||
COPY package.json package-lock.json ./
|
||||
RUN npm ci --omit=dev
|
||||
COPY . .
|
||||
|
||||
FROM node:22-slim
|
||||
WORKDIR /app
|
||||
COPY --from=build /app .
|
||||
RUN npx --yes playwright install --with-deps chromium
|
||||
ENV NODE_ENV=production
|
||||
# De SQLite-database hoort op het persistente volume te staan; zonder deze
|
||||
# default zou een vergeten DB_PATH stilletjes naar de container-fs schrijven
|
||||
# en bij elke redeploy alle data verliezen.
|
||||
ENV DB_PATH=/data/korting.db
|
||||
VOLUME /data
|
||||
EXPOSE 3000
|
||||
CMD ["node", "server.js"]
|
||||
```
|
||||
|
||||
- [ ] **Step 2: Update `README.md`**
|
||||
|
||||
Add one short paragraph after the existing "Hoe het werkt" section explaining the fallback (read the current `README.md` first to place it consistently with the existing tone/structure):
|
||||
|
||||
```markdown
|
||||
Als het ophalen van de prijs op de gewone manier mislukt (netwerkfout, of
|
||||
geen prijs gevonden in de HTML), probeert de app het één keer opnieuw met
|
||||
een echte headless-browser (Playwright/Chromium) — nuttig voor pagina's
|
||||
die de prijs pas met JavaScript renderen, of die eenvoudige HTTP-clients
|
||||
blokkeren. Geen garantie tegen elke vorm van bot-detectie, maar dekt een
|
||||
reëel deel van de gevallen die de gewone aanpak mist.
|
||||
```
|
||||
|
||||
- [ ] **Step 3: Manual build/run verification (this is the first task that needs a real browser)**
|
||||
|
||||
Run:
|
||||
```bash
|
||||
docker build -t korting .
|
||||
```
|
||||
Expected: build succeeds; watch for the `playwright install --with-deps chromium` step completing without error (it will print apt package installation output — that's expected).
|
||||
|
||||
Run:
|
||||
```bash
|
||||
docker run --rm -d --name korting-fallback-smoketest -p 3002:3000 \
|
||||
-e BASIC_AUTH_USER=admin -e BASIC_AUTH_PASS=test \
|
||||
-e SMTP_HOST=mail.sjoerd.app -e SMTP_PORT=587 \
|
||||
-e SMTP_FROM=korting@sjoerd.app -e NOTIFY_EMAIL=test@example.com \
|
||||
-e DB_PATH=/data/korting.db -v korting-fallback-smoketest-data:/data \
|
||||
korting
|
||||
sleep 2
|
||||
curl -u admin:test -X POST http://localhost:3002/products \
|
||||
-d "name=Fallback test" -d "url=https://example.com"
|
||||
curl -u admin:test -X POST http://localhost:3002/check-now
|
||||
docker logs korting-fallback-smoketest 2>&1 | tail -20
|
||||
```
|
||||
Expected: the `check-now` request succeeds (302); the container logs show the `[check] 1 product(en) gecontroleerd: ...` summary line. This URL (`https://example.com`) has no parseable price, so the primary fetch will return `not_found`, which **does** trigger the fallback — confirm the logs show no crash and no unhandled Playwright error (a `not_found` after the fallback also finds nothing is the expected, correct outcome here — the goal of this check is proving the browser actually launches and runs inside the container without error, not that this particular blank test page has a price).
|
||||
|
||||
Clean up:
|
||||
```bash
|
||||
docker stop korting-fallback-smoketest
|
||||
docker volume rm korting-fallback-smoketest-data
|
||||
docker rmi korting
|
||||
```
|
||||
|
||||
- [ ] **Step 4: Commit**
|
||||
|
||||
```bash
|
||||
git add Dockerfile README.md
|
||||
git commit -m "chore: install Chromium in the Docker image for the fallback scraper"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## After this plan
|
||||
|
||||
Deploy the merged result to Coolify (redeploy the existing `korting.sjoerd.app` application — same app, no new resources needed) via the `deploying-to-coolify` skill, and verify against the actual reported failing URL
|
||||
(`https://www.iciparisxl.nl/chanel/bleu-de-chanel/parfum-verstuiver/p/BP_967866`)
|
||||
using Coolify's deployment/application logs — not a direct request from this
|
||||
machine, since the user has since rotated the app's basic-auth credentials
|
||||
and testing now happens on their side.
|
||||
Reference in New Issue
Block a user