
Connects Claude directly to your Lastest visual regression testing instance, letting you trigger test runs, review screenshot diffs, approve or reject baselines, and manage test suites through conversation. You can kick off builds for specific branches, inspect failure details, and update test configurations without leaving your chat context. Useful when you're iterating on UI changes and want to check for regressions inline, or when triaging CI failures and need to approve intentional visual changes on the spot. Works with both self-hosted and cloud Lastest deployments.
Free, open-source visual regression testing with AI-generated tests
Record it. Test it. Ship it.
Website • Wiki • Features • Quick Start • How It Works • Why Lastest • Comparison • Commands • Config
Visual regression testing is either expensive, flaky, or painful to maintain.
Meanwhile, you just need to know: "Did my last commit break the UI?"
Lastest is a free, self-hosted visual regression testing platform that records your tests, writes them with AI, runs them anywhere, and fixes them when they break — all in one tool.
1. Point it at your app
2. Record your user flows (point-and-click, no code)
3. AI generates resilient test code with multi-selector fallback
4. Every test runs inside an Embedded Browser container (EB stack required)
5. Screenshots compared with 3 diff engines (pixelmatch, SSIM, Butteraugli)
6. Review and approve visual changes — or let AI auto-classify them
When self-hosted, your data stays on your server and your screenshots never leave your infra.
Lastest adapts to how you want to build tests — from fully manual to fully autonomous.
Open the recorder, click through your app, hit stop. Lastest captures every interaction and generates deterministic Playwright code — no AI involved, no API keys needed. You own the test code and can edit it by hand.
Best for: Teams that don't want AI, air-gapped environments, simple flows.
AI generates, fixes, or enhances tests — but you review and approve before anything is saved. Feed it a URL and get a test back. Import OpenAPI specs or user stories and AI extracts test cases. When a test breaks, AI proposes a fix and you decide whether to accept it.
Best for: Day-to-day development, iterating on tests, fixing breakages fast.
One click kicks off an 11-step pipeline: check settings, select repo, set up environment, scan routes & apply testing template, plan functional areas, review plan, generate tests, run them, fix failures (up to 3 attempts per test), re-run, and report results. Uses specialized sub-agents (Orchestrator, Planner, Scout, Diver, Generator, Healer). The agent pauses and asks for help only when it hits something it can't resolve on its own. You resume and it picks up where it left off.
Best for: Onboarding a new project, generating full coverage from scratch, CI bootstrapping.
Every test runs inside an Embedded Browser (EB) pod — a containerized Chromium with live CDP streaming back to the UI. There is no local-Playwright fallback and no remote-runner daemon anymore: the host never launches its own browser, and if an EB is unavailable a test fails rather than falling back. The EB stack is therefore required for every run and recording, even in local development.
lastest-eb container yourself (docker run … -p 9223:9223 -p 9224:9224) and register it in Settings → Runners to get a token. This is the replacement for old "remote runners": you bring a browser on the machine/network/OS you want, not a Playwright daemon.EB pods handle both running and recording. Builds can be triggered manually (click Run), by webhook (PR opened/updated), from CI/CD (GitHub Action or the @lastest/runner CLI, which only triggers a build and polls — execution still happens in the EB pool), or on a schedule (cron-based automation). Smart Run analyzes git diffs to run only affected tests.
Migrating from remote runners? The standalone runner daemon that executed Playwright on remote machines has been retired. To run tests on your own machine/network, register a BYO Embedded Browser instead. The
@lastest/runnernpm package lives on as a lightweight, browser-free CI trigger client only.
Tests are recorded or generated once, then stored as code. Every subsequent run re-executes the same code, captures new screenshots, and diffs them against approved baselines.
Create tests (one-time) Run tests (forever)
┌──────────────────────┐ ┌──────────────────────┐
│ Manual recording │ │ Execute Playwright │
│ — or — │ ────▶ │ Capture screenshots │
│ AI generation │ save │ Diff against baseline │
│ — or — │ │ Review changes │
│ Play Agent autonomy │ │ Approve/reject │
└──────────────────────┘ └──────────────────────┘
AI may be used here No AI needed here
/run and /builds pages were retired into it, bringing the build-history graph (header drawer), the smart/all/comparison run split-button, "Accept all safe" on unsorted diffs, live EB streaming during a build, and the step-label editor.key, label, baseUrl, releaseLabel, refreshedAt), with per-environment variables and baselines. Run one suite against UAT and PROD without the two fighting over one set of approvals; promote baselines between environments and survive a vendor sandbox refresh. Fully additive — a repo that never creates an environment behaves exactly as before.credentials parameter, never through variable substitution — so rotating a password does not change codeHash, invalidate a baseline, or land in plaintext run records.method / url / headers / query / body / auth + assertions). Runs in-process with no browser or EB dispatch, feeds the same step-comparison and verdict pipeline as browser tests, and can be generated from captured network calls. Includes a burst/load runner for firing concurrent requests at an endpoint (single up-front SSRF validation, per-connection re-check).expect() calls passed/failed with expected vs actual values, error messages, and code line references./r/<slug> report — no login required for viewers. Includes optional AI-written demo notes, session video (WebM with MP4 conversion + fallback), zip download of artifacts, and social-share flows with media-aware cards for X, YouTube, and TikTok. Links are revokable and reuse a stable URL on re-publish.GET /export + POST /import).lastest_api_*) for the MCP server, VS Code extension, CI scripts, and cross-instance migration. Revokable per-user with labels.suggest_app_fix tool.validate_diff / decide_diff), scoping analysis to a single diff for precise, low-noise verdicts./agents is the single entry point for agent work: one roster of every agent working the selected repo, what each is doing, what it is blocked on, and which embedded browsers it is holding. Paused (browser released) and blocked-on-a-human (browser still held) are shown as different states, and escalations from parked sessions and pushed-back tasks are merged into one queue. Drill into a row to reach the QA agent, the Triage agent, the Healer, or the Explorer (Ranger, Play and QuickStart are one-shot onboarding flows and have no roster row). Pro-plan gated.test_maintenance, flaky_test); real regressions, environment issues, unclassified failures and cases a reviewer already decided are left red with the reason recorded. Hard stops: a per-test attempt budget (default 2, counted from ai_fix versions since the test last passed or was hand-edited), a per-build cap (default 5), an identical-error no-progress guard, one campaign per repo at a time, and heal/verify timeouts. Every patch is a versioned ai_fix edit you can revert. /healer-agent owns the toggle, the budgets, a Stop button and a per-test outcome ledger.Date.now() and new Date() with fixed values for deterministic screenshots.Math.random() for consistent outputs.lastest-eb containers you host and register yourself. JPEG streaming with configurable quality/framerate, WebSocket auth, concurrent contexts. If no EB is available the run fails — the host never launches its own browser.@lastest/runner) — Lightweight, browser-free npm client that triggers a build on your Lastest server and polls for results (writes GITHUB_OUTPUT / GITHUB_STEP_SUMMARY, exits non-zero on failure or --fail-on-changes). Execution runs server-side in the EB pool; the old distributed-execution runner daemon has been retired.core/ kernel (contracts, browser, data, jobs, storage) and 20+ self-contained plugins/ (explorer, qa-agent, app-map, recorder, share, ci, scheduling, gamification, rca, api-test, design-system, a11y, data-sources, …) that talk to core only through capability contracts, plus pure libs/ packages. Enforced by an architecture test (pnpm arch).@lastest/mcp-server) exposing a consolidated, resource-oriented surface of 29 tools for AI agent integration: run/verify tests, review and decide diffs, approve baselines, create/heal tests, suggest app fixes, publish shares, check coverage. Install via npx @lastest/mcp-server./api/mcp is an OAuth-protected resource: RFC 8414 metadata, RFC 7591 dynamic client registration, PKCE S256, and RFC 9728 WWW-Authenticate on 401, so an agent platform can connect knowing nothing but the URL. A tool-access policy narrows the surface by caller — read (observe), write (create/update/run/approve), full (deletes and anything that makes data public, API keys only) — by rewriting tool schemas rather than rejecting calls after the fact.document.modelContext (Chrome origin trial, ChatGPT desktop/Work, Codex "site tools"), with a polyfill fallback, a consent dialog, a team-level toggle, and public-share tools. Backed by a cookie-authed /api/mcp/session sibling endpoint that reuses the same server, narrowed to 16 agent-facing tools with route ids taken from the page, never from the agent.packages/ocr-service, pnpm ocr:up) rather than in-process; set OCR_SERVICE_URL to enable OCR selectors and text-region-aware diffing./api/v1/) for IDE integration.window.__APP_STATE__, Redux stores, etc.) for complex assertions.Running tests requires the Embedded Browser stack — there is no local-Playwright fallback. Bring up the database, the host app, and the k3d EB cluster together:
git clone https://github.com/las-team/lastest.git
cd lastest
docker compose up -d # postgres on :5432 (named volume `lastest-pgdata`)
pnpm install
# Configure env (required before pnpm dev — without EB_PROVISIONER the app
# starts but no test can ever provision an Embedded Browser).
cp .env.example .env.local
cat >> .env.local <<EOF
EB_PROVISIONER=kubernetes
EB_NAMESPACE=lastest
EB_IMAGE=lastest-embedded-browser:latest
LASTEST_URL=http://host.k3d.internal:3000
SYSTEM_EB_TOKEN=$(openssl rand -hex 32)
EOF
pnpm db:push # apply schema
pnpm stack # REQUIRED: create k3d cluster + build/import EB image
pnpm dev # http://localhost:3000
Open http://localhost:3000.
docker compose down (data persists in the lastest-pgdata volume).docker compose down -v.The dev app runs on the host while EB pods are dynamically provisioned into a local k3d cluster — one EB per test. Without pnpm stack running, no test can execute or record.
pnpm stack # create k3d cluster, build + import the EB image
pnpm stack:status # cluster + EB jobs + host /api/health
pnpm stack:logs # tail EB pod logs
pnpm stack:refresh # rebuild the EB image after editing packages/embedded-browser
pnpm stack:stop # delete the cluster
The .env.local keys the EB stack requires (already added by the Quick Start block above):
EB_PROVISIONER=kubernetes
EB_NAMESPACE=lastest
EB_IMAGE=lastest-embedded-browser:latest
LASTEST_URL=http://host.k3d.internal:3000
SYSTEM_EB_TOKEN=<openssl rand -hex 32>
DATABASE_URL=postgresql://lastest:lastest@localhost:5432/lastest
Note on defaults:
EB_PROVISIONERdefaults to'none'in code, so omitting it letspnpm devstart cleanly but every test will silently fail to provision an EB. Always set it for any environment that runs tests.
See k8s/ and scripts/k3d-*.sh for the manifests and bootstrap scripts.
k3d ≥ 5.6, kubectl, openssl (the EB stack — no local-Playwright fallback)┌──────────────────┐ ┌─────────────┐ ┌─────────────┐
│ Create Tests │ ──▶ │ Run │ ──▶ │ Review │
│ │ │ │ │ │
│ Record manually │ │ Embedded │ │ Approve/ │
│ AI-assisted │ │ Browser or │ │ Reject │
│ Play Agent auto │ │ remote/CI │ │ changes │
└──────────────────┘ └─────────────┘ └─────────────┘
One-time cost No AI per run New baseline
(AI optional) (pure Playwright) saved
Create: Build tests your way — record manually in the browser, let AI generate from a URL or spec, or let the Play Agent autonomously scan your entire app.
Run: Every test executes inside an Embedded Browser pod — provisioned on demand, one per test, and streamed live to the UI. Trigger from the UI, a webhook, a schedule, or CI/CD via the @lastest/runner trigger CLI. Screenshots are captured at key steps. No AI needed — pure Playwright execution at zero cost. There is no local-Playwright fallback and no remote-runner daemon; the EB stack is required.
Compare: New screenshots are diffed against baselines using your chosen engine (pixelmatch, SSIM, or Butteraugli). Text-region-aware comparison available. Accessibility audits run automatically.
Review: Visual diffs are classified (unchanged/flaky/changed). AI can optionally auto-classify with confidence scores. Approve intentional changes — they become the new baseline.
Fix: When tests break, AI can propose fixes (human-in-the-loop) or the Play Agent can fix and re-run autonomously. When a failure is a real regression, the Fix-the-App Advisor can instead suggest a fix to your application code (never auto-applied). Or edit the code by hand — your choice.
End-to-end view of how Lastest fits into a CI/CD workflow — from git push through preview-env validation to merge approval.
All test execution happens in Embedded Browser pods. When a build starts, the Lastest server provisions EB pods on demand into the Kubernetes cluster (k3d locally, your cluster in production) — one browser per test, bounded by a per-build worker pool with warm-pool keep-alive. Each pod runs Playwright, streams its screen live over CDP, captures screenshots, and reports results back. Pod egress is restricted and pod creation is throttled (CNI burst protection). The host process never launches a browser; if no EB is available, the run fails rather than falling back.
For CI/CD, the @lastest/runner CLI creates a build over HTTP and polls for the result — it carries no browser and executes nothing locally; the EB pool does all the work server-side.
| Capability | Lastest | Percy | Applitools | Chromatic | Argos | Meticulous | Playwright |
|---|---|---|---|---|---|---|---|
| Price | Free self-hosted / hosted plans | Paid | Paid | Paid | Paid | Paid | Free |
| Screenshot volume (self-hosted) | Unlimited | Limited | OSS only | Limited | Limited | None | Unlimited |
| Self-hosted | Yes | No | Enterprise | No | OSS core | No | Yes |
| Open source | FSL-1.1-ALv2 | SDKs only | SDKs only | Storybook | MIT core | No | Apache-2.0 |
| No-code recording | Yes | No | Low-code | No | No | Session | Codegen |
| AI test generation | Yes | No | NLP | No | No | Session-based | No |
| AI auto-fix tests | Yes | No | No | No | No | Auto-maintain | No |
| Autonomous agent | Yes (Play Agent) | No | No | No | No | No | No |
| AI diff analysis | Yes | AI Review Agent | Visual AI | No | No | Deterministic | No |
| Multi-engine diffing | 3 engines | No | Visual AI | No | No | No | No |
| Text-region-aware diffing | Yes | No | No | No | No | No | No |
| Spec-driven test gen | Yes | No | No | No | No | No | No |
| Approval workflow | Yes | Yes | Yes | Yes | Yes | PR-based | No |
| Accessibility | axe-core | No | No | Enterprise | ARIA snaps | No | No |
| Route discovery | Yes | No | No | No | No | No | No |
| Multi-tenancy | Yes | Projects | Enterprise | Projects | Teams | Projects | No |
| Figma integration | Yes | No | Yes | No | No | No | No |
| Google Sheets data | Yes | No | No | No | No | No | No |
| Debug mode | Yes | No | No | No | Traces | No | Trace |
| CI trigger CLI | Yes (@lastest/runner) | Cloud | Cloud | Cloud | Cloud | Cloud | No |
| Embedded browser execution | Yes (container + live stream) | No | No | No | No | No | No |
| Headless API testing | Yes | No | No | No | No | No | No |
| Public share links | Yes (watermarked /r/) | No | No | No | No | No | No |
| AI app-code fix advisor | Yes | No | No | No | No | No | No |
| Local AI (Ollama) | Yes | No | No | No | No | No | No |
| Cross-OS consistency | 12 stabilization features | No | No | No | Stabilization engine | No | No |
| GitHub Action | Yes | Cloud-only | Cloud-only | Cloud-only | Cloud-only | Cloud-only | No |
| GitLab integration | Yes (self-hosted) | Yes | Yes | No | No | No | No |
| Test composition | Yes | No | No | No | No | No | No |
| Testing templates | 8 presets | No | No | No | No | No | No |
| Setup/teardown orchestration | Yes | No | No | No | No | No | No |
| Branch baseline management | Yes | Yes | Yes | Yes | No | No | No |
| Scheduled test runs | Yes (cron) | Cloud | Cloud | Cloud | Cloud | Cloud | No |
| MCP server (AI agent API) | Yes (29 tools) | No | No | No | No | No | No |
| Remote MCP with OAuth 2.1 | Yes (DCR + scoped tools) | No | No | No | No | No | No |
| WebMCP (browser agent tools) | Yes | No | No | No | No | No | No |
| App map / swarm crawler | Yes (multi-EB) | No | No | No | No | Crawler | No |
| Environments (PROD/UAT) | Yes (per-env baselines) | No | Enterprise | No | No | No | No |
| Per-repo secret store | Yes (encrypted) | Cloud | Cloud | Cloud | Cloud | Cloud | No |
| WCAG compliance scoring | Yes (0–100) | No | No | No | No | No | No |
| AI failure triage | Yes | No | No | No | No | No | No |
| Assertion tracking | Yes | No | No | No | No | No | No |
| Agent monitoring | Yes (real-time SSE) | No | No | No | No | No | No |
| In-app bug reports | Yes (auto-context) | No | No | No | No | No | No |
| Gamification | Yes (leaderboard + achievements) | No | No | No | No | No | No |
| Cross-instance migration | Yes (API export/import) | No | No | No | No | No | No |
| API tokens | Yes (long-lived Bearer) | Cloud | Cloud | Cloud | Cloud | Cloud | No |
@lastest/runner CLI/r/ report with AI demo notes, session video, and X/YouTube/TikTok social cardspnpm dev # Start development server on localhost:3000
pnpm build # Production build
pnpm start # Start production server
pnpm lint # Run ESLint
pnpm test # Run unit tests (Vitest)
pnpm test:watch # Run unit tests in watch mode
pnpm test:coverage # Run tests with coverage report
pnpm test:ui # Run tests with Vitest UI
pnpm db:studio # Open Drizzle Studio for database inspection
pnpm db:push # Push schema changes to database
pnpm db:generate # Generate Drizzle migrations
pnpm db:reset # Reset database (drops all tables + removes screenshots/baselines)
pnpm db:seed # Seed test data
pnpm test:visual # Run visual tests via CLI (see below)
pnpm test:integration # Integration tests (Vitest, needs a database)
pnpm arch # Architecture test — enforce the core/plugin boundary
# OCR service container (required for OCR selectors + text-region-aware diffing)
pnpm ocr:up # docker compose up -d --build ocr (set OCR_SERVICE_URL)
pnpm ocr:down
# Local k3d cluster — hosts dynamically-provisioned EB Job pods (no app, no db)
pnpm stack # create cluster + build/import EB image
pnpm stack:refresh # rebuild + import EB image (alias of stack:refresh:eb)
pnpm stack:refresh:eb # same
pnpm stack:status # cluster + EB jobs/pods + host /api/health
pnpm stack:logs # tail EB pod logs
pnpm stack:stop # delete cluster
In-depth docs for every integration live on the Lastest Wiki. The README keeps things at a glance — click through for flags, payloads, and CI examples.
| Guide | What it covers | Wiki |
|---|---|---|
| CLI Test Runner (CI/CD) | pnpm test:visual --repo-id <id> for GitHub Actions / other pipelines; auto-captures GITHUB_HEAD_REF / GITHUB_REF_NAME / GITHUB_SHA | CI/CD Integration |
| GitHub Action | Reusable composite action las-team/lastest/action@main — zero local Playwright, triggers a build on your Lastest server (executed in the EB pool); outputs status + build URL + counts | CI/CD Integration · GitHub Integration |
| Smart Run | Diff-based test selection — only tests affected by changed files run, comparing the feature branch against the default branch via GitHub/GitLab API | Running Tests |
| Self-Hosted Deployment | pnpm deploy:zima (ZimaBoard / CasaOS via docker compose) and pnpm deploy:olares (Olares via kubectl); shared multi-stage Dockerfile, GET /api/health. Required env: POSTGRES_PASSWORD, BETTER_AUTH_SECRET, SYSTEM_EB_TOKEN | Docker Deployment |
| CI Trigger CLI | @lastest/runner on npm — lastest-runner trigger -r <repo> -t <token> -s <url> creates a build and polls for results (no local browser); execution runs server-side in the Embedded Browser pool | CI Trigger CLI |
| MCP Server | npx @lastest/mcp-server --url <…> --api-key <…> exposes 29 consolidated tools (run/verify/heal/approve/decide-diff/suggest-app-fix/publish-share/coverage/…) for Claude and other agents; structured JSON responses | MCP Server |
| Remote MCP / WebMCP | /api/mcp over OAuth 2.1 (dynamic client registration, PKCE, read/write/full tool policy) for agent platforms; document.modelContext tools for in-browser agents, behind a consent dialog and a team toggle | MCP Server |
| Environments & Connectors | PROD / UAT / prerelease environments with per-environment variables and baselines; Veeva Vault and Salesforce connectors bound to an environment and a credential set | Settings Reference |
| QA Agent | Eight-phase autonomous suite builder — preflight, discovery (static route scan + live EB crawl), plan, human plan review, generate, run, heal, summary | Agent Monitoring |
| Scheduled Runs | Cron-based automated builds with presets (daily 3am, weekly, hourly, every 15min) or custom expressions; auto-disable after 5 consecutive failures | Scheduled Runs |
| Google Sheets Integration | Spreadsheet-backed test data — per-team OAuth, multi-tab spreadsheets, custom header row, fixed ranges; surfaces values on the test Vars tab | Google Sheets |
| Custom Webhooks | POST build.completed payloads (status / counts / git refs / build URL) to any HTTP endpoint, with custom method + headers | Custom Webhooks |
| VSCode Extension | lastest-vscode — Test Explorer in the Activity Bar, run tests from the editor, live status bar, real-time WebSocket updates; powered by /api/v1/ REST + SSE | VSCode Extension API |
| API Tokens | Long-lived Bearer tokens for programmatic access (CI runners, MCP server, REST clients) | API Tokens |
| Bug Reports | In-app reporting with auto-captured browser/network/console context, optional GitHub-issue creation | Bug Reports |
| Agent Monitoring | Real-time SSE activity feed tracking Play Agent sessions step-by-step | Agent Monitoring |
| Gamification | "Beat the Bot" — scoring, seasons, leaderboards, Bug Blitz multiplier events | Gamification |
| Test Migration | Cross-instance export / import of tests, areas, and configs via REST API or in-app UI | Test Migration |
All configuration lives under a unified Settings page. Per-section deep dives live on the Settings Reference wiki — the table below is the quick map.