tsproxy/README.md

197 lines
9.6 KiB
Markdown
Raw Normal View History

2026-05-11 15:12:13 +00:00
# tsproxy
A tiny TCP forwarder that joins your tailnet via [tsnet](https://pkg.go.dev/tailscale.com/tsnet)
and exposes one or more tailnet services as local ports — without running
`tailscaled` on the host.
Useful when you want to reach a single service on your tailnet from a machine
you'd rather not fully enroll (work laptop, ephemeral VM, etc). The proxy
itself becomes a tailnet node, so it can be tightly scoped with tags and ACLs.
## Install
```sh
go install github.com/smw/tsproxy@latest # or build locally:
go build -o /usr/local/bin/tsproxy .
```
## Quick start
1. (Optional) In your tailnet policy, add a tag for the proxy and an ACL that limits what it can reach. Example (adapt to your hosts):
```json
"tagOwners": {
"tag:tsproxy": ["you@example.com"]
},
"acls": [
{ "action": "accept", "src": ["autogroup:members"], "dst": ["*:*"] },
{ "action": "accept", "src": ["tag:tsproxy"], "dst": ["webhost:80", "dbhost:5432"] }
]
```
2. Mint a one-shot auth key at
<https://login.tailscale.com/admin/settings/keys>:
**not reusable**, **not ephemeral**, tagged `tag:tsproxy`. The key is only
needed for the first run — tsnet persists its own credentials after that.
3. First run:
```sh
TS_AUTHKEY=tskey-auth-... tsproxy \
--name tsproxy \
--forward 127.0.0.1:8080=webhost:80 \
--forward 127.0.0.1:5432=dbhost:5432
```
4. Subsequent runs don't need `TS_AUTHKEY`:
```sh
tsproxy -f 127.0.0.1:8080=webhost:80 -f 127.0.0.1:5432=dbhost:5432
```
Then `curl http://127.0.0.1:8080` hits `webhost:80` on your tailnet.
## Flags
| Flag | Short | Default | Description |
|---|---|---|---|
| `--forward` | `-f` | *(required, repeatable)* | Forward rule `LOCAL=TARGET`, e.g. `127.0.0.1:8080=myhost:80`. |
| `--name` | `-n` | `tsproxy` | Hostname advertised on the tailnet. |
| `--dir` | | `~/.config/tsproxy/<name>` | State directory (node identity, keys). |
| `--verbose` | `-v` | `false` | Verbose tsnet logging. |
2026-08-01 01:12:04 +00:00
| `--probe-interval` | | `5s` | How often a *down* target is re-probed. No probing happens while it is up. |
2026-08-01 01:26:44 +00:00
| `--probe-timeout` | | `10s` | How long a reachability probe may take before the target counts as down. |
| `--dial-timeout` | | `30s` | How long a forwarded connection may take to reach the target before it is failed. |
2026-08-01 01:12:04 +00:00
| `--idle-timeout` | | `5m` | Close a forwarded connection after this long with no traffic either way (`0` disables). |
| `--trace` | | `false` | Log every accept, dial, probe and teardown, for diagnosing stalls. |
2026-05-11 15:12:13 +00:00
Target can be any MagicDNS name, short hostname, or tailnet IP.
2026-08-01 01:37:46 +00:00
### Prefer a tailnet IP for the target
The dial path resolves a target in two stages. A literal IP is used as-is with
no lookup at all. A *name* is checked against the in-memory tailnet map, and on
a miss falls through to a real **system DNS** query — which can stall for
seconds on a short hostname, entirely on the local machine, before any packet
goes near the target.
That shows up as forwarded dials alternating between milliseconds (resolver
warm) and seconds (resolver cold), which is long enough for a client to hit its
own timeout while it sits connected and unserved. If you see `SLOW DIAL` in the
logs, try the tailnet IP (`100.x.y.z:PORT`) or the fully-qualified MagicDNS name
instead of the short one.
### Tuning for a low-latency link
The defaults are conservative because a tailnet dial may need to renegotiate a
path or fall back to DERP. If your target is a millisecond away, much tighter
values are reasonable and will surface a dead target in about a second:
```sh
tsproxy -f 127.0.0.1:8080=100.x.y.z:80 --dial-timeout 2s --probe-timeout 1s
```
**Fix resolution before tightening these.** A slow dial caused by system DNS is
not a slow link, and a timeout short enough to cut it off turns connections that
would have worked into failures — each one resetting a client and closing the
port until the next probe, which then pays the same DNS cost. Confirm dials are
consistently fast (`--trace` shows `connected to target in ...`) and only then
lower the budgets.
2026-05-11 15:12:13 +00:00
## Running under launchd (macOS)
An example LaunchAgent plist is included as
[dev.tsproxy.plist](dev.tsproxy.plist). Install it per-user:
```sh
cp dev.tsproxy.plist ~/Library/LaunchAgents/
launchctl load ~/Library/LaunchAgents/dev.tsproxy.plist
```
Manage it:
```sh
launchctl list | grep tsproxy
launchctl kickstart -k gui/$(id -u)/dev.tsproxy # restart after edits
launchctl unload ~/Library/LaunchAgents/dev.tsproxy.plist
tail -f /tmp/tsproxy.log
```
On first run, add the auth key via an `EnvironmentVariables` block in the
plist, then remove it once the state directory has been populated.
## Notes
- **One node, many forwards.** All `--forward` rules share a single tsnet
identity, so you get one device in the admin console and one ACL subject.
2026-07-31 21:32:27 +00:00
- **Startup is strict.** Every local address is bound once at startup as a
check; if any fails, the process exits — partial success is confusing under
a supervisor.
2026-08-01 01:12:04 +00:00
- **The local port tracks the target.** A forward only keeps its local port
bound while the target is believed reachable. So an unavailable target means
`connection refused` on the local port, not a connect that immediately EOFs,
and clients back off the way they would against a genuinely down service.
- **Probing only happens while down.** One probe runs at startup to decide
whether to bind at all. After that, real connections are the health signal
and nothing polls — a healthy target is never dialled except to carry
traffic, so services that log every connection stay quiet. Probing resumes
(every `--probe-interval`) only once something has failed, and stops again as
soon as the target answers. The cost of not polling is that a target which
dies unnoticed leaves the port bound until something tries to use it: that
one connection is accepted and then reset, and the port closes behind it.
2026-08-01 00:29:50 +00:00
- **One failed connection closes the port.** A client that reconnects the
instant its connection breaks would otherwise beat the next probe and be
accepted into a forward with nothing behind it. So any connection that fails
against the target unbinds the port immediately, and it stays unbound until a
probe says the target is back. That probe is scheduled straight away, so a
one-off failure against a healthy target costs a few milliseconds of refusal,
not a whole interval. Other live connections are left alone — one failure
stops new work but isn't enough to declare the target dead.
2026-08-01 01:26:44 +00:00
- **Dials are bounded, but generously.** A tailnet peer that is routable but
dead neither accepts nor refuses, so an unbounded dial parks forever holding
a local socket that nothing can reclaim. Forwarded dials are therefore capped
at `--dial-timeout`, which is deliberately much larger than `--probe-timeout`:
a probe is a health check that costs nothing to retry and can afford to be
impatient, while a forwarded dial has a client waiting on it and would rather
wait than fail. Tailnet dial latency is very uneven — tens of milliseconds on
a warm path, seconds when the path has to be renegotiated or fall back to
DERP — so a probe-sized budget here fails connections to a healthy target.
A dial taking more than a quarter of its budget is logged without `--trace`,
since that is the shape of trouble before it becomes a failure.
2026-07-31 21:32:27 +00:00
- **Target loss drops live connections.** A tailnet peer can vanish without the
userspace TCP stack ever erroring on an established connection, which leaves
2026-08-01 01:12:04 +00:00
local sockets hanging and apps waiting on a dead link. Two things catch this.
A connection that fails outright takes the port down immediately, and the
probe that follows resets everything still in flight. A connection that
merely goes silent is caught by `--idle-timeout`, since with polling switched
off there is nothing else watching it.
- **Idle connections are closed.** Once `--idle-timeout` passes with no bytes
moving *either way*, the connection is reset. Traffic in one direction counts,
so a long upload or download is never reaped mid-stream. This is what bounds
the case where the target accepts a connection and then goes silent — but it
cannot tell that apart from a connection legitimately sitting idle, so a
shell session or a database pool left quiet past the timeout is dropped too
and has to reconnect. Lower it to notice dead targets sooner, raise it if
your clients hold connections open across long gaps, `0` to disable. Reaping
an idle connection never closes the local port: an unused connection says
nothing about whether the target is healthy.
2026-07-31 21:54:09 +00:00
- **Failures are reset, not closed.** How a connection ends is forwarded
faithfully. A target that closes cleanly gives the local client a FIN (an
ordinary EOF); a target that resets, errors, or disappears gives it an RST.
This matters for clients that hold idle connections: a FIN leaves the socket
writable, so a pooled client's *next* write still succeeds and only the one
after it fails, whereas an RST fails the very next read or write. It also
keeps a truncated response from looking like a complete one. The trade-off
is that an RST discards whatever was still queued in the send buffer — a
stream torn down this way was already incomplete, so flagging it beats
delivering a partial result that looks whole.
2026-07-31 21:32:27 +00:00
- **Probes are real connections.** Each probe opens and immediately closes a
TCP connection to the target. Chatty services may log these; raise
`--probe-interval` to quiet them down, at the cost of slower detection.
2026-05-11 15:12:13 +00:00
- **State directory matters.** Losing `~/.config/tsproxy/<name>/` means the
node re-registers on next launch and will need a fresh `TS_AUTHKEY`.
## License
[MIT](LICENSE)