196 lines
9.6 KiB
Markdown
196 lines
9.6 KiB
Markdown
# tsproxy
|
|
|
|
A tiny TCP forwarder that joins your tailnet via [tsnet](https://pkg.go.dev/tailscale.com/tsnet)
|
|
and exposes one or more tailnet services as local ports — without running
|
|
`tailscaled` on the host.
|
|
|
|
Useful when you want to reach a single service on your tailnet from a machine
|
|
you'd rather not fully enroll (work laptop, ephemeral VM, etc). The proxy
|
|
itself becomes a tailnet node, so it can be tightly scoped with tags and ACLs.
|
|
|
|
## Install
|
|
|
|
```sh
|
|
go install github.com/smw/tsproxy@latest # or build locally:
|
|
go build -o /usr/local/bin/tsproxy .
|
|
```
|
|
|
|
## Quick start
|
|
|
|
1. (Optional) In your tailnet policy, add a tag for the proxy and an ACL that limits what it can reach. Example (adapt to your hosts):
|
|
|
|
```json
|
|
"tagOwners": {
|
|
"tag:tsproxy": ["you@example.com"]
|
|
},
|
|
"acls": [
|
|
{ "action": "accept", "src": ["autogroup:members"], "dst": ["*:*"] },
|
|
{ "action": "accept", "src": ["tag:tsproxy"], "dst": ["webhost:80", "dbhost:5432"] }
|
|
]
|
|
```
|
|
|
|
2. Mint a one-shot auth key at
|
|
<https://login.tailscale.com/admin/settings/keys>:
|
|
**not reusable**, **not ephemeral**, tagged `tag:tsproxy`. The key is only
|
|
needed for the first run — tsnet persists its own credentials after that.
|
|
|
|
3. First run:
|
|
|
|
```sh
|
|
TS_AUTHKEY=tskey-auth-... tsproxy \
|
|
--name tsproxy \
|
|
--forward 127.0.0.1:8080=webhost:80 \
|
|
--forward 127.0.0.1:5432=dbhost:5432
|
|
```
|
|
|
|
4. Subsequent runs don't need `TS_AUTHKEY`:
|
|
|
|
```sh
|
|
tsproxy -f 127.0.0.1:8080=webhost:80 -f 127.0.0.1:5432=dbhost:5432
|
|
```
|
|
|
|
Then `curl http://127.0.0.1:8080` hits `webhost:80` on your tailnet.
|
|
|
|
## Flags
|
|
|
|
| Flag | Short | Default | Description |
|
|
|---|---|---|---|
|
|
| `--forward` | `-f` | *(required, repeatable)* | Forward rule `LOCAL=TARGET`, e.g. `127.0.0.1:8080=myhost:80`. |
|
|
| `--name` | `-n` | `tsproxy` | Hostname advertised on the tailnet. |
|
|
| `--dir` | | `~/.config/tsproxy/<name>` | State directory (node identity, keys). |
|
|
| `--verbose` | `-v` | `false` | Verbose tsnet logging. |
|
|
| `--probe-interval` | | `5s` | How often a *down* target is re-probed. No probing happens while it is up. |
|
|
| `--probe-timeout` | | `10s` | How long a reachability probe may take before the target counts as down. |
|
|
| `--dial-timeout` | | `30s` | How long a forwarded connection may take to reach the target before it is failed. |
|
|
| `--idle-timeout` | | `5m` | Close a forwarded connection after this long with no traffic either way (`0` disables). |
|
|
| `--trace` | | `false` | Log every accept, dial, probe and teardown, for diagnosing stalls. |
|
|
|
|
Target can be any MagicDNS name, short hostname, or tailnet IP.
|
|
|
|
### Prefer a tailnet IP for the target
|
|
|
|
The dial path resolves a target in two stages. A literal IP is used as-is with
|
|
no lookup at all. A *name* is checked against the in-memory tailnet map, and on
|
|
a miss falls through to a real **system DNS** query — which can stall for
|
|
seconds on a short hostname, entirely on the local machine, before any packet
|
|
goes near the target.
|
|
|
|
That shows up as forwarded dials alternating between milliseconds (resolver
|
|
warm) and seconds (resolver cold), which is long enough for a client to hit its
|
|
own timeout while it sits connected and unserved. If you see `SLOW DIAL` in the
|
|
logs, try the tailnet IP (`100.x.y.z:PORT`) or the fully-qualified MagicDNS name
|
|
instead of the short one.
|
|
|
|
### Tuning for a low-latency link
|
|
|
|
The defaults are conservative because a tailnet dial may need to renegotiate a
|
|
path or fall back to DERP. If your target is a millisecond away, much tighter
|
|
values are reasonable and will surface a dead target in about a second:
|
|
|
|
```sh
|
|
tsproxy -f 127.0.0.1:8080=100.x.y.z:80 --dial-timeout 2s --probe-timeout 1s
|
|
```
|
|
|
|
**Fix resolution before tightening these.** A slow dial caused by system DNS is
|
|
not a slow link, and a timeout short enough to cut it off turns connections that
|
|
would have worked into failures — each one resetting a client and closing the
|
|
port until the next probe, which then pays the same DNS cost. Confirm dials are
|
|
consistently fast (`--trace` shows `connected to target in ...`) and only then
|
|
lower the budgets.
|
|
|
|
## Running under launchd (macOS)
|
|
|
|
An example LaunchAgent plist is included as
|
|
[dev.tsproxy.plist](dev.tsproxy.plist). Install it per-user:
|
|
|
|
```sh
|
|
cp dev.tsproxy.plist ~/Library/LaunchAgents/
|
|
launchctl load ~/Library/LaunchAgents/dev.tsproxy.plist
|
|
```
|
|
|
|
Manage it:
|
|
|
|
```sh
|
|
launchctl list | grep tsproxy
|
|
launchctl kickstart -k gui/$(id -u)/dev.tsproxy # restart after edits
|
|
launchctl unload ~/Library/LaunchAgents/dev.tsproxy.plist
|
|
tail -f /tmp/tsproxy.log
|
|
```
|
|
|
|
On first run, add the auth key via an `EnvironmentVariables` block in the
|
|
plist, then remove it once the state directory has been populated.
|
|
|
|
## Notes
|
|
|
|
- **One node, many forwards.** All `--forward` rules share a single tsnet
|
|
identity, so you get one device in the admin console and one ACL subject.
|
|
- **Startup is strict.** Every local address is bound once at startup as a
|
|
check; if any fails, the process exits — partial success is confusing under
|
|
a supervisor.
|
|
- **The local port tracks the target.** A forward only keeps its local port
|
|
bound while the target is believed reachable. So an unavailable target means
|
|
`connection refused` on the local port, not a connect that immediately EOFs,
|
|
and clients back off the way they would against a genuinely down service.
|
|
- **Probing only happens while down.** One probe runs at startup to decide
|
|
whether to bind at all. After that, real connections are the health signal
|
|
and nothing polls — a healthy target is never dialled except to carry
|
|
traffic, so services that log every connection stay quiet. Probing resumes
|
|
(every `--probe-interval`) only once something has failed, and stops again as
|
|
soon as the target answers. The cost of not polling is that a target which
|
|
dies unnoticed leaves the port bound until something tries to use it: that
|
|
one connection is accepted and then reset, and the port closes behind it.
|
|
- **One failed connection closes the port.** A client that reconnects the
|
|
instant its connection breaks would otherwise beat the next probe and be
|
|
accepted into a forward with nothing behind it. So any connection that fails
|
|
against the target unbinds the port immediately, and it stays unbound until a
|
|
probe says the target is back. That probe is scheduled straight away, so a
|
|
one-off failure against a healthy target costs a few milliseconds of refusal,
|
|
not a whole interval. Other live connections are left alone — one failure
|
|
stops new work but isn't enough to declare the target dead.
|
|
- **Dials are bounded, but generously.** A tailnet peer that is routable but
|
|
dead neither accepts nor refuses, so an unbounded dial parks forever holding
|
|
a local socket that nothing can reclaim. Forwarded dials are therefore capped
|
|
at `--dial-timeout`, which is deliberately much larger than `--probe-timeout`:
|
|
a probe is a health check that costs nothing to retry and can afford to be
|
|
impatient, while a forwarded dial has a client waiting on it and would rather
|
|
wait than fail. Tailnet dial latency is very uneven — tens of milliseconds on
|
|
a warm path, seconds when the path has to be renegotiated or fall back to
|
|
DERP — so a probe-sized budget here fails connections to a healthy target.
|
|
A dial taking more than a quarter of its budget is logged without `--trace`,
|
|
since that is the shape of trouble before it becomes a failure.
|
|
- **Target loss drops live connections.** A tailnet peer can vanish without the
|
|
userspace TCP stack ever erroring on an established connection, which leaves
|
|
local sockets hanging and apps waiting on a dead link. Two things catch this.
|
|
A connection that fails outright takes the port down immediately, and the
|
|
probe that follows resets everything still in flight. A connection that
|
|
merely goes silent is caught by `--idle-timeout`, since with polling switched
|
|
off there is nothing else watching it.
|
|
- **Idle connections are closed.** Once `--idle-timeout` passes with no bytes
|
|
moving *either way*, the connection is reset. Traffic in one direction counts,
|
|
so a long upload or download is never reaped mid-stream. This is what bounds
|
|
the case where the target accepts a connection and then goes silent — but it
|
|
cannot tell that apart from a connection legitimately sitting idle, so a
|
|
shell session or a database pool left quiet past the timeout is dropped too
|
|
and has to reconnect. Lower it to notice dead targets sooner, raise it if
|
|
your clients hold connections open across long gaps, `0` to disable. Reaping
|
|
an idle connection never closes the local port: an unused connection says
|
|
nothing about whether the target is healthy.
|
|
- **Failures are reset, not closed.** How a connection ends is forwarded
|
|
faithfully. A target that closes cleanly gives the local client a FIN (an
|
|
ordinary EOF); a target that resets, errors, or disappears gives it an RST.
|
|
This matters for clients that hold idle connections: a FIN leaves the socket
|
|
writable, so a pooled client's *next* write still succeeds and only the one
|
|
after it fails, whereas an RST fails the very next read or write. It also
|
|
keeps a truncated response from looking like a complete one. The trade-off
|
|
is that an RST discards whatever was still queued in the send buffer — a
|
|
stream torn down this way was already incomplete, so flagging it beats
|
|
delivering a partial result that looks whole.
|
|
- **Probes are real connections.** Each probe opens and immediately closes a
|
|
TCP connection to the target. Chatty services may log these; raise
|
|
`--probe-interval` to quiet them down, at the cost of slower detection.
|
|
- **State directory matters.** Losing `~/.config/tsproxy/<name>/` means the
|
|
node re-registers on next launch and will need a fresh `TS_AUTHKEY`.
|
|
|
|
## License
|
|
|
|
[MIT](LICENSE)
|