← All articles

MCPProxy v0.61.0: Reconnect Storms, Scans That Actually Run, and 47 UX Fixes

TL;DR

v0.61.0 is out — 25 commits over v0.60.0. Three things worth your attention:

  1. Reconnect storms are gone. A dead upstream used to get re-dialed every 30 seconds forever — about 9,000 requests a day, and a separate 5-second token scan added roughly 8,600 more. Both are now on real backoff.
  2. The free security scan actually runs. Telemetry said 12 of 157 capable installs had ever scanned a server. The scan is now automatic on every newly added server, plus a one-time sweep of what you already have.
  3. Secrets are masked in the Web UI. The activity drawer used to print the credential it had just flagged, in cleartext, next to a Copy button.

Plus 47 UX findings from two audits, and a batch of smaller fixes. Details below.

The Reconnect Storm

This one started as a real incident. A third-party MCP vendor’s alarms went off twice because of us, and our account got paused.

The cause: MCPProxy’s supervisor reconciles every 30 seconds. For every enabled server that is not connected, it planned a connect. It did that regardless of failure history. The managed client underneath had exponential backoff and a give-up counter, and the reconcile path walked straight past both. One attempt is about three HTTP requests. Every 30 seconds, forever, is roughly 9,000 requests a day against a server that had already told us “401”.

The fix landed in two rounds, because the first round was not enough.

Round one (#1014) made the supervisor consult the client’s retry policy before planning a connect. Backoff is the usual ladder — 1s, 2s, 4s, capped at 5 minutes.

Round two (#1039) found three holes in round one:

  • OAuth failures were not counted. SetOAuthError bumps an OAuth retry counter, never the general one. The new gate only read the general one, so a 401 read as “no failures yet” and got re-dialed on every tick — the exact behaviour we set out to stop, for the error class most likely to cause it. OAuth failures now follow their own ladder: 5min, 15min, 1h, 4h, 24h.
  • The “waiting for login” state was unreachable. A server that needs OAuth should park until you log in, because re-dialing cannot succeed and each attempt costs a real request. Two separate bugs prevented that: the deferred-OAuth error was wrapped twice on its way out so the type check never matched, and the code parked the server and then immediately called SetError, which unconditionally moved it back to the error state. The park survived microseconds. It parks for real now, and the UI says Sign-in required instead of reporting the server as healthy.
  • The token scan was a tighter storm than the one we fixed. A background scan runs every 5 seconds looking for newly persisted OAuth tokens, and it fired a reconnect whenever any token existed for a failing server. A failing server usually still holds the expired token it just failed with — so it redialed roughly 8,600 times a day on its own, out-pacing the supervisor loop. It now fingerprints the token it last retried with (a truncated hash over access + refresh + expiry + write stamp, never the raw token) and only fires when a login or refresh writes a different one.

The fourth change went the other way. “Gave up” used to mean never again. The retry ladder exhausts in about an hour, so a laptop sleep, a VPN drop, or overnight upstream maintenance left the server silently dead until a human noticed. A given-up server is now probed once every 30 minutes: about 48 requests a day instead of 2,880, and it heals itself.

Separately, #1042 added Retry-After support for 429 responses. It is captured at the transport layer — a RoundTripper reads the header off the raw response — because by the time a rate-limit error reaches the MCP client layer, the header is gone.

Manual reconnects, login flows, reconnect_on_use, and config-change reconnects are all unaffected. Only the automatic redial is gated.

Docs: OAuth authentication.

Security Scanning That Actually Runs

MCPProxy ships a free, offline, in-process Tool Poisoning Attack scan. It needs no Docker and no network. Almost nobody was running it, and that was our fault, not theirs.

The only automatic trigger required trust_mode: "scan" (never the default) and the server to be quarantined and never scanned and to have no approval baseline. Nothing ever rescanned an existing server. Telemetry: 12 of 157 capable installs had ever scanned anything.

v0.61.0 adds two informational paths (#1031):

  • Every newly added, enabled server gets one baseline scan, regardless of trust mode. Informational — it does not gate anything, it just tells you what is in the tool descriptions you just admitted.
  • A one-shot sweep on upgrade. On first start after upgrading, servers you already have that have never been scanned get swept through the same path. It is backgrounded, delayed 45 seconds so upstreams can connect, serialized, and cancellable. It never delays startup. If every candidate fails because the servers are still connecting, the “done” marker is not written and the next start retries.

The kill switch is security.auto_baseline_scan (default on) or MCPPROXY_AUTO_BASELINE_SCAN.

Scanning is also reachable now from the surfaces people use (#1032). quarantine_security is the most-used security MCP tool — 64 installs against 12 for the dedicated scanning tools — and it never mentioned scanning at all. It gains two operations:

  • scan_server — runs the offline baseline scan in-process, waits up to 8 seconds for a verdict, otherwise returns the job id.
  • get_scan_report — latest verdict, risk score, counts, and up to 10 findings.

list_quarantined and inspect_tools now carry a one-line scan status, so an unscanned server reads as “never scanned — run scan_server first” rather than looking clean.

Two false Docker gates are gone. “Scan All Servers” did not render on a fresh install because the overview counted only persisted scanners and missed the built-in in-process one. It was also disabled without Docker. Docker is only needed for the optional deep scanners; the baseline is built in. The CLI help said the same wrong thing and has been rewritten.

Docs: tool scanner, security quarantine.

The Drawer Printed the Secret

This one is embarrassing enough to spell out. A tool call flagged critical · aws_access_key, and a few hundred pixels below the red detection panel the Web UI rendered the arguments — including AKIAIOSFODNN7EXAMPLE in cleartext, with a Copy button.

The raw value was in the API response, so the fix is server-side (#1051). Redacting in the browser was rejected: the value would still be on the wire and in the network log.

Every value the detector recognises inside arguments, responses, and error messages is now replaced with a preview — AKIAIOSFODNN7EXAMPLE becomes AKIA…****. Enough to know which credential to rotate, not enough to use one. Private keys go whole, envelope and body.

A cross-model review round found the API boundary necessary but not sufficient, and every finding was reproduced against the code before being fixed:

  • /events streamed the raw payload to every SSE subscriber — Web UI, tray, mcpproxy activity watch — and the frontend logged it into the DevTools console. SSE events are emitted at completion time, before the async detector has a verdict, so there is nothing to gate on: the SSE writer masks unconditionally.
  • Error messages were unmasked. Upstream errors routinely quote the payload they choked on, so masking only the response just moved the leak.
  • Metadata was unmasked, and the intent reason is prose the calling agent wrote. An agent explaining “rotating AKIA…” put the value straight back.
  • The tool-call endpoints are a separate store with no detection metadata and bypassed masking entirely.
  • A private-key block ended at the first -----END, so a payload quoting an EC key inside an RSA key left the outer key’s remainder readable.

Full values stay reachable through the one surface that already gates them: mcpproxy activity export --include-bodies. That is the incident-response path, not a browsing one. The drawer badges both payload panels Masked and says where the full values live.

Docs: sensitive data detection.

47 UX Findings

We ran two audits against v0.60.0: every routed view of the Web UI in light and dark, first-run and populated, at 1440px, 820px and 390px, with programmatic WCAG contrast sampling; and a settings-parity pass over the macOS tray. Together they produced 52 ranked findings. v0.61.0 closes 47.

The ones you will notice:

Numbers that agree with each other. Activity, Usage and the Dashboard each counted calls their own way, so the same hour produced three different totals. Calls are counted once now, in one place (#1048, #1052). If a number looks wrong now, it is wrong — that is an improvement.

A real first run. A fresh install landed you on a page that had nothing to say. There is now a call to action, an onboarding path, and a “what needs me” panel on the landing page (#1049).

Honest server states. Connection state displays and the connect modal no longer report states they have not verified (#1053).

Contrast, theme, mobile. WCAG AA contrast fixes throughout, a proper system-theme option alongside light and dark, and a mobile layout that works (#1054).

Repeated calls grouped. An agent retrying the same tool twelve times is one row now, not twelve. Folding is on by default, remembered per browser, and switched off whenever the table is sorted by anything other than time — adjacency only means “repeated” in time order. A run never folds calls with different outcomes together.

Tray parity. All 16 findings from the macOS tray audit (#1055, #1056). The menu-bar icon badges server health with a severity dot, and the button carries a real accessibility label — “MCPProxy — 13/29 servers, 942 tools — 3 server errors” — so VoiceOver reads the state instead of the word “MCPProxy”. Enable, disable, restart and login no longer swallow errors. The Servers submenu groups by state and folds disabled servers away. There is a native Tools item with BM25 search: the old tool search called an endpoint that has no q parameter and read a field it never returns, so it had always returned an empty list. And code_execution_pool_size is correctly marked restart-required.

/search is gone as a separate route. It was a third search surface, routed but not in the sidebar, duplicating the header box and the Tools page. It redirects to Tools now, carrying the query and hash so old links keep their state.

Docs: dashboard, activity log.

Under the Hood

  • Telemetry heartbeat v9 (#1029). The old counters only saw scan jobs, which most installs never start, so the two detection paths that actually run for ordinary users — the tool-change gate and the prompt-poisoning filter — emitted nothing and the fleet read as “the scanner never runs”. v9 adds counters for both, plus a trust-mode distribution. Counts of invocations only: no server names, no tool names, no verdicts, no matched checks. The whole sub-object is still dropped when every counter is zero, so an install that scans nothing sends the same shape it did before. Telemetry is opt-out with MCPPROXY_TELEMETRY=false; what it sends is documented in full.
  • Graceful shutdown flushes the final heartbeat (#1037). It used to be dropped on exit.
  • Scanner data races fixed (#1038) — live pointers shared across the scanner registry and engine.
  • Tray --listen works (#1015, #1036). mcpproxy-tray serve --listen 0.0.0.0:8181 silently dropped the flag: the tray never parsed its own command line, so the core came up on the config default and anything probing the advertised port concluded the daemon was down. The follow-up caught a worse one — normalizeListen("127.0.0.1:8181") returned ":8181", quietly widening a loopback bind the user had pinned to all interfaces. A pinned host is now preserved verbatim, a bare port defaults to loopback, and an explicit all-interfaces request is honoured as written. A pinned listen that hits a port conflict now fails loudly instead of silently moving to another port your clients are not configured for.

Upgrading

brew upgrade smart-mcp-proxy/mcpproxy/mcpproxy
# or grab the signed DMG / installer / .deb / .rpm from the release page

Installers for every platform are on the release page. macOS builds are signed and notarized; the tray updates itself.

Nothing in this release changes config format or on-disk state. If you have OAuth upstreams that have been failing quietly, expect them to move to Sign-in required on first start instead of retrying — that is the intended behaviour, and it is the first time you will actually see them.

Install docs: getting started.


Further reading