I’ve been doing remote attestation work on confidential VMs. This post is one piece of that: whatever the attestation ends up looking like, the machine still needs an identity it can hold and use afterwards, and in my case that means authenticating to other services with mutual TLS.
An application that does mutual TLS authenticates with a client certificate,
and on disk that is almost always two files. client.crt is the one you show:
it carries a public key and a CA’s signature over that key, and it isn’t a
secret. client.key is the one you sign the handshake with, and the server
checks that signature against the public key in the certificate you just showed
it.
So the identity is entirely in the second file, and your process reads it into memory at startup. From that moment the key is in a heap dump, a core file, a swapped page, a hypervisor snapshot of the VM’s memory, and in reach of anything that gets code execution in the process. Whoever copies it can act as that machine everywhere it is trusted: the secret store, the internal API, the database that maps certificate subjects to roles.
A key in a file is not really a machine identity. It’s a bearer credential that happens to be stored on a machine, and whoever reads it becomes that machine, anywhere, until somebody notices and revokes the certificate.
The threat model is narrow. What I’m defending against is an attacker who gets read access inside the guest: a file read, a core dump, an SSRF that reaches the filesystem, a backup that went somewhere it shouldn’t. I want them to leave without a working copy of the machine’s identity.
What I’m not defending against is someone with persistent code execution using that identity while they’re still on the box. That’s a different problem and a TPM doesn’t solve it.
So the key I need has four properties:
- It signs TLS handshakes, because that is how the identity gets used.
- It never exists outside the machine, so reading memory or a disk gets an attacker nothing they can carry away.
- It needs no second credential to reach it, because whatever that credential was would immediately become the new thing worth stealing.
- It works with an ordinary TLS stack, because I don’t want a bespoke protocol between two services that already speak TLS.
A TPM does all four. It will sign with a key and it will not hand the key over, so the process never holds a secret at all. On the machine I’ve been using that’s a virtual TPM, and I wrote about the chip itself in an earlier post. Being a confidential VM helps with exactly one item on that exposure list, since guest memory is encrypted against the host and a snapshot taken from underneath gets ciphertext. Everything else on the list happens inside the guest, where memory encryption doesn’t help.
why not something else
Three alternatives come up before a TPM does, and all three are reasonable. Each of them misses at least one of the four properties.
File permissions are worth tightening, and they still fail the second
property. SELinux, RLIMIT_CORE=0 and mlockall close some vectors, but a
process that uses the key has the key in its address space, and ptrace,
/proc/pid/mem and code execution in that process are all untouched by the
mode bits on a file that has already been read.
Short-lived certificates fail the second property too, and partly the third. They shrink the window a stolen key is useful in, which is worth more than anything in this post if you can only do one thing. They don’t stop the copy, and you still need some credential to authenticate the renewal.
A KMS is the real competitor. Cloud KMS or Vault’s transit engine holds a key that never leaves, so it passes the second property, and for a lot of teams it is the right answer. What it fails is the third. A KMS has to authenticate you somehow, so there is now a credential sitting on the box that unlocks the signing key, and that credential is the thing an attacker copies instead. That moves the problem instead of solving it, unless that credential is itself bound to hardware.
There’s a practical cost on top of that, which is that every cold handshake now depends on a network round trip to a service that can be down. I would rather not put a remote dependency in the path that decides whether two of my services can talk to each other.
So it comes down to the third property (no second credential to reach the key). Something has to be the root: held by the machine, impossible to copy out, and reachable without presenting anything first. A TPM is that, and everything else here (a KMS token, a short-lived workload certificate, a PKCS#11 token you’d have to go and buy) can hang off it rather than compete with it.
None of this makes a TPM more secure than a well-run KMS. It’s local, it needs no credential to reach, and on most machines built in the last decade it is already there.
what TLS asks of a key
A TPM will sign with a key and do nothing else with it. That is also all TLS needs.
In TLS 1.3 a peer proves it holds the private key for the certificate it just
sent by signing a message defined in
CertificateVerify. That
message is 64 bytes of 0x20, the string TLS 1.3, client CertificateVerify,
a zero byte, and then the transcript hash, which is a running hash of every
handshake message the two sides have exchanged so far. All of it gets hashed,
and that digest is what the key signs. A server authenticating with a
certificate signs the same structure with server in place of client, which
is what stops either side’s signature being replayed back at the other.
Two things follow from the transcript being in there. Because it includes fresh randomness from both peers, the signature is good for this one connection and no other, so it can’t be replayed. And because there’s one transcript per handshake, the key is asked for exactly one signature over one digest.
The private key never decrypts anything and never derives the session keys.
Those come out of an ephemeral Diffie-Hellman exchange it takes no part in.
Once CertificateVerify is on the wire, the key is done for the life of the
connection.
Go asks for the same thing. A tls.Certificate holds a PrivateKey, and what
that field has to satisfy is
crypto.Signer:
type Signer interface {
Public() crypto.PublicKey
Sign(rand io.Reader, digest []byte, opts crypto.SignerOpts) ([]byte, error)
}
crypto/tls builds the CertificateVerify message itself, hashes it itself, and
calls Sign with the digest. It never asks for key material. PKCS#11 modules
and cloud KMS keys plug into Go through the same interface, and a TPM fits it
just as well.
TPM2_Sign has the same shape: digest in, signature out, key stays put. So the
library I wrote is a Sign that forwards to the TPM, plus some care around
attaching to the right key.
the client
Here’s the whole client. Nothing in it names the key:
leaf, err := x509.ParseCertificate(certDER)
if err != nil {
return err
}
// The certificate picks the key.
key, err := tpmtls.OpenForCertificate(tpmtls.DefaultDevice, leaf)
if err != nil {
return err
}
defer key.Close()
conn, err := tls.Dial("tcp", "store.example.com:443", &tls.Config{
MinVersion: tls.VersionTLS13,
Certificates: []tls.Certificate{key.TLSCertificate(leaf.Raw)},
})
There’s no handle or key path, and no device configuration beyond a default.
The tls.Config doesn’t mention a TPM either. It’s the same config you’d
write for a key in a file, and it works just as well in an http.Transport, in
tls.Listen, or on an http.Server.
OpenForCertificate comes from
go-tpm-tls, a thin layer I
wrote over go-tpm and
go-tpm-tools. It takes the
certificate and nothing else, which I went back and forth on, so I’ll explain
it.
A key in a TPM sits at a handle, which is a number like 0x81000004 saying
where the object lives. The obvious thing is to put that number in
configuration, and I wrote it that way first, which was a mistake. The handle
is chosen by the process that provisioned the key, so a value that works on one
machine fails on the next, and it fails at the first handshake rather than at
startup, which is a bad time to find out. But your application already holds
the certificate it’s about to present, and the public key inside that
certificate says exactly which key in the TPM to sign with. So the lookup reads
the persistent handles, compares public keys, and picks the match. If nothing
matches you get tpmtls.ErrNotFound at startup.
It also matches the lifetimes of the two halves. The key stays in the TPM for
the life of the machine and the certificate over it is what rotates, so the
certificate is what your application gets handed fresh and the key is what it
has to go and find. If you do know the handle, tpmtls.Open takes one.
Two smaller things in that snippet. tpmtls.DefaultDevice is /dev/tpmrm0,
the kernel resource manager, and you want that rather than the raw /dev/tpm0
because it gives every open file descriptor its own context, so processes
sharing the TPM don’t evict each other’s objects. That per-descriptor detail
matters again in the benchmarks. Opening it needs root or membership of the
tss group.
And you should pin MinVersion: tls.VersionTLS13 rather than inheriting Go’s
default of 1.2. The usual reason for that is mostly obsolete: in TLS 1.2 a
server certificate’s key can be asked to decrypt a premaster secret under the
static RSA key exchange suites, and a TPM signing key cannot do that. Modern Go
already disables those suites, and client authentication in TLS 1.2 is a
signature either way, so this is not a real risk anymore. Pin it anyway,
because TLS 1.3 makes signing the only thing that can be asked of the key, and
because the numbers later in this post were measured on it.
what the code assumes
The code above depends on six things being true, and only one of them is under your control. The other five were decided by another process before yours started, and your code just attaches to the result.
A signing key already exists in the TPM. Something else put it there,
typically an attestation agent that generated the key inside the TPM, bound the
public half into hardware evidence, and had it certified. go-tpm-tls has no way
to create a key, deliberately. It also won’t evict one. Close lets go of the
key and leaves the handle alone, because detaching from a key you didn’t
provision shouldn’t destroy it for everyone else on the machine.
The key is unrestricted. A restricted key, which is what an attestation key
is, will only sign data the TPM itself hashed or produced. TPM2_Sign wants a
validation ticket proving that, and a transcript hash handed in by crypto/tls
has no such ticket, so the TPM answers TPM_RC_TICKET. If your agent
provisioned an attestation key and a TLS key, you want the second one.
The key is ECDSA, on P-256 or P-384. RSA keys load and sign, but not the
way crypto/tls asks, which I’ll come back to.
The key has no auth value. go-tpm-tls loads with a nil session, which resolves to a null session, so a key guarded by a password or a policy will not load. That is a limitation, not a security choice: an auth value would be one more thing an attacker on the box has to obtain. If your provisioning sets one, this package won’t work for you yet.
The key sits at a persistent handle. Persistent handles run from
0x81000000 to 0x81FFFFFF, survive a reboot, and can be reached by whatever
process comes along later. A transient handle works too, but only over the TPM
connection that created it, and it costs about ten times as much per signature.
If you have any say in provisioning, ask for persistent.
Your process can open /dev/tpmrm0. This is the only one you can fix
without re-provisioning the machine. It needs root or the tss group.
Since you inherit the other five, check what you actually got rather than
trusting the process that provisioned it. NonExportable reads the attributes
back out of the TPM:
ok, err := key.NonExportable()
if err != nil {
return err
}
if !ok {
return fmt.Errorf("key at %#x can be duplicated out of the TPM", key.Handle())
}
That’s the fixedTPM attribute, which is the TPM telling you it will refuse to
duplicate the private key to another TPM. Paired with sensitiveDataOrigin at
creation time, which says the TPM generated the key rather than being handed
one, you know the key has never existed anywhere else. NonExportable only
reads the first of those two back; the second is a property of how it was
provisioned.
From a shell it’s tpm2_getcap and tpm2_readpublic, worth running once so
you know what the good case looks like. This is a GCE SEV-SNP guest with
tpm2-tools 5.6:
$ sudo tpm2_getcap handles-persistent
- 0x81000004
$ sudo tpm2_readpublic -c 0x81000004
name: 000ba090006025e535419d16b12e7bdd219d20077aa4bd0fc99bc41cc5ed643005b3
name-alg:
value: sha256
attributes:
value: fixedtpm|fixedparent|sensitivedataorigin|userwithauth|sign
raw: 0x40072
type:
value: ecc
curve-id:
value: NIST p256
scheme:
value: ecdsa
scheme-halg:
value: sha256
fixedtpm and sensitivedataorigin are the two that matter, sign without
restricted is the third, and there is no private half anywhere in that output
because there is no way to ask for one.
making a key to try this with
On a real machine an attestation agent creates the key, but you need one to run any of the above, so here’s the short version.
A warning first. This makes a key that cannot leave the TPM, and that is all it does. Nothing binds it to what the machine booted, and no verifier anywhere has evidence of where it came from, so a certificate issued over it attests to nothing beyond “some TPM held this”. That is fine for trying the code in this post, but it is not provisioning. A real attestation agent creates the key, binds the public half into a quote, and has it certified against that evidence, which is what makes the key worth trusting.
In the snippet, client is go-tpm-tools/client:
// A throwaway key for trying out the code in this post.
template := tpm2.Public{
Type: tpm2.AlgECC,
NameAlg: tpm2.AlgSHA256,
Attributes: tpm2.FlagSign | tpm2.FlagSensitiveDataOrigin |
tpm2.FlagUserWithAuth | tpm2.FlagFixedTPM | tpm2.FlagFixedParent,
ECCParameters: &tpm2.ECCParams{
Sign: &tpm2.SigScheme{Alg: tpm2.AlgECDSA, Hash: tpm2.AlgSHA256},
CurveID: tpm2.CurveNISTP256,
},
}
const handle = tpmtls.Handle(0x81000004)
rwc, err := tpm2.OpenTPM(tpmtls.DefaultDevice)
if err != nil {
return err
}
defer rwc.Close()
// Creates the key inside the TPM and persists it at the handle.
created, err := client.NewCachedKey(rwc, tpm2.HandleOwner, template, handle)
if err != nil {
return err
}
created.Close() // let go of it; the key stays at the handle
That template is five of the six assumptions written out in code.
FlagSensitiveDataOrigin means the TPM generates the private key rather than
being handed one. FlagFixedTPM and FlagFixedParent mean it can’t be
duplicated out. FlagSign with no FlagRestricted means it will sign a digest
you give it, which is what CertificateVerify needs. The curve is P-256,
NewCachedKey persists it rather than leaving it transient, and since no auth
value is supplied the key ends up with an empty one, which is what lets
go-tpm-tls load it with a null session.
You still need something to present, and key.CertificateRequest produces a
CSR signed by the TPM that you can hand to whatever CA you’re testing against.
Evict the key when you’re done. Persistent slots are few and a key left behind survives reboots:
tpm2.EvictControl(rwc, "", tpm2.HandleOwner, handle, handle)
Owner auth is the empty string on the confidential VMs I ran this on. That is not true everywhere.
the other end doesn’t know
Here’s the verifying end. There’s no TPM on this side and nothing here knows about one:
pool := x509.NewCertPool()
pool.AppendCertsFromPEM(caPEM)
srv := &http.Server{
Addr: ":8443",
TLSConfig: &tls.Config{
MinVersion: tls.VersionTLS13,
ClientAuth: tls.RequireAndVerifyClientCert,
ClientCAs: pool,
},
}
The server checks a signature against the public key in a certificate that chains to a CA it trusts. Where the private half was while that signature was made is invisible to it, which means you can roll this out one service at a time, from either direction.
None of this is client-specific. The server’s own key is a crypto.Signer too,
so putting it in the TPM is the same swap and the same call:
key, err := tpmtls.OpenForCertificate(tpmtls.DefaultDevice, serverLeaf)
if err != nil {
return err
}
defer key.Close()
srv.TLSConfig.Certificates = []tls.Certificate{key.TLSCertificate(serverLeaf.Raw)}
A server signs one CertificateVerify per full handshake, same as a client, so everything below about cost applies, and the throughput ceiling matters a lot more on a listener than on a dialler. Its clients need no changes, because from their side this is still an ordinary certificate.
That is also the limit. A TPM changes how strong the possession claim is, not what the far end verifies. If you want the verifier to know that the key is in a TPM, and which machine’s TPM, that has to be in the certificate, which means it has to come from attestation at issuance time, because the handshake carries none of it.
four ways it breaks
Four things broke while I was getting this working. None of the errors mention TLS or the TPM, which made them annoying to track down, so I’ve listed them by the message you’ll see.
salt length must be rsa.PSSSaltLengthAuto
You’re using an RSA key. TLS 1.3
requires RSA-PSS with a
salt as long as the digest, so crypto/tls asks for exactly that. A TPM picks
its own salt length and you can’t override it, so go-tpm-tools refuses the
request rather than produce a signature the peer would reject.
The key loads fine, so this fails at signing time in the middle of a handshake. Use P-256 or P-384 and the problem goes away.
key at 0x81000004 cannot sign: restricted keys are not supported
You attached to a restricted key, almost certainly the attestation key. It signs only what the TPM itself hashed, which it proves with a validation ticket, and a TLS transcript hash is caller-supplied and has no ticket.
You can reproduce this without go-tpm-tls. Hand a restricted key 32 random bytes and ask it to sign them, which is what CertificateVerify does:
$ sudo tpm2_sign -c restricted.ctx -d -g sha256 -s ecdsa -o sig.bin digest.bin
ERROR: Esys_Sign(0x3E0) - tpm:parameter(3):invalid ticket
ERROR: Unable to run tpm2_sign
The same command against an unrestricted key signs those bytes.
TPM_RC_TICKET is the refusal that go-tpm-tools is reporting to you when it
says restricted keys are not supported.
This one fails at load, and that’s deliberate. The check runs when you attach rather than when you handshake, so a misconfigured key takes down startup instead of one connection an hour later.
tpmtls: no persistent key matches
That’s ErrNotFound, and it means the certificate you’re holding and the
contents of the TPM disagree. Either the attestation agent hasn’t provisioned a
key yet, or it provisioned a different one than the certificate is over, which
is what a stale certificate looks like after re-provisioning. Test for it with
errors.Is, then list the handles and compare.
remote error: tls: error decrypting message, on the first read
This one is TLS rather than TPM, and it’s the one that caught me.
In TLS 1.3 the client sends its Certificate, its CertificateVerify and its
Finished in a single flight, and it does not wait to hear whether any of that
was accepted. tls.Dial returns you a working connection either way. If the
server refuses your signature it sends an alert, and you get that alert on the
next read.
So a workload whose identity the far end won’t accept looks exactly like one that connected. If you health check by dialling, you’ve checked that TCP works and that the server’s certificate is valid, and nothing about your own. Read a byte:
conn, err := tls.Dial("tcp", addr, cfg)
if err != nil {
return err
}
defer conn.Close()
// The server's verdict on our certificate arrives as an alert on the next
// read, not as a dial error. Without this the probe passes either way.
if _, err := conn.Write(probe); err != nil {
return err
}
if _, err := conn.Read(make([]byte, 1)); err != nil {
return err
}
If you’re going through net/http you get this for free, since any real
request reads a response. The trap is a readiness probe that only dials.
what it costs
You pay for the TPM once per full handshake, not per request.
A TPM signs in milliseconds where software signs in microseconds. I measured
how much of that reaches a connection in
go-tpm-tls-bench, running
on Google Cloud Confidential VMs where the vTPM is implemented in hypervisor
software. What follows is an n2d-standard-2 with AMD SEV-SNP and a P-256 key,
single-threaded, measured against a software key of the same curve:
| scenario | key | median | p95 | throughput | signatures |
|---|---|---|---|---|---|
| raw signing, 100 signatures | TPM, persistent | 2.21 ms | 2.61 ms | 452 sig/s | 100 |
| software | 0.06 ms | 0.08 ms | 16591 sig/s | 100 | |
| TPM, transient | 20.44 ms | 21.20 ms | 49 sig/s | 100 | |
| 50 connections, no resumption | TPM | 3.35 ms | 3.83 ms | 295 conn/s | 50 |
| software | 1.38 ms | 1.51 ms | 707 conn/s | 50 | |
| 50 connections, resumption | TPM | 0.94 ms | 1.59 ms | 856 conn/s | 1, 49 resumed |
| software | 0.95 ms | 1.13 ms | 1011 conn/s | 1, 49 resumed | |
| 1 connection, 50 requests | TPM | 0.02 ms | 0.03 ms | 40043 req/s | 1 |
| software | 0.02 ms | 0.02 ms | 39133 req/s | 1 |
The signature is most of a cold handshake, 2.21 ms out of 3.35 ms. A TPM runs one command at a time and signing is serialized behind a lock for that reason, so a machine tops out at a few hundred new mutually authenticated connections per second. That limit is per machine, since each machine has its own TPM, so it doesn’t get better by putting more replicas behind one of them.
In the last two scenarios the two keys land within noise of each other. The signature counts explain why: 50 signatures for 50 fresh connections, 1 for 50 resumed ones, and 1 for 50 requests over a reused connection. The TPM didn’t get faster, there were just fewer handshakes.
That ceiling is also a DoS vector. On a listener anyone can reach, forcing full handshakes is a cheap way to burn a few hundred signatures a second, and since the lock has no timeout, every handshake in the process queues behind it. Don’t put a TPM key on a public listener. Use it on an outbound client, on an internal listener, or behind something that limits handshake rate.
Two results surprised me, and I only found the second because I measured it wrong the first time.
The transient row is a factor of ten, and fifteen on Intel TDX. The kernel resource manager context-saves transient objects between commands, so every signature pays to swap the key back into the TPM, while a persistent object just sits in the TPM’s own storage. I expected that to matter, but not by a factor of ten.
What I didn’t expect is that the cost is per file descriptor rather than per
object. A persistent key that nothing else is touching signs in 2.08 ms, and
the same key signs in 20.04 ms while an unrelated transient object is loaded on
the same file descriptor. Same factor of ten, on a key that was never
transient. On a separate descriptor the same key costs 2.10 ms. So a persistent
key is necessary but not sufficient, and what you control is what else your
process loads on the descriptor it signs on, which should be nothing. That is
also why I’d use Open rather than New: the key owns its descriptor and
nothing else in the process can put a transient object on it.
Curve choice only costs you on cold handshakes, and on SEV-SNP it costs less than I expected. P-384 signs in 2.37 ms against P-256’s 2.21 ms on this machine, which is 226 cold handshakes a second against 295. I wouldn’t quote that gap to two significant figures: an earlier run on the same instance type put it at 16% rather than 7%, so the premium is small and noisy on this platform. On Intel TDX the same comparison came out at 65%. Resumption erases the difference either way, which matters if a policy like CNSA 1.0 puts you on P-384.
The P-256 table above is a single run on one n2d-standard-2, and the TDX
figures come from a c3-standard-4. Both have a hypervisor-implemented vTPM,
and none of it says anything about discrete TPM hardware.
what it proves
A handshake signed this way proves the connection comes from the machine whose TPM holds the key, and that it happened now rather than being a replay, since the transcript carries fresh randomness from both sides. There are three things it does not prove.
It says nothing about the code running on that machine. An attacker inside your
process can sign for as long as they’re in it, and so can anything else on the
box, because the key has no auth value and the access control on it is
therefore just whether you can open /dev/tpmrm0, which means root or the
tss group. The key is bound to the machine, not to your workload. Narrowing
it to one process is a separate problem and the TPM doesn’t solve it.
It doesn’t keep that attacker out of your traffic either. The session keys are in ordinary memory. What stays in the TPM is the long-term identity, not the bytes on the connection.
And it depends where the vTPM itself lives, which is a platform choice and not
something confidential computing settles for you. SEV-SNP encrypts guest memory
against the host; it says nothing about who implements the TPM. On the machine
I measured these numbers on, Google’s, the vTPM is hypervisor software, so the
key is out of reach of everything in the guest but not out of reach of the
platform. You can see this in dmesg: SEV: SNP running at VMPL0. means the
kernel holds the most privileged level itself, so nothing is running beneath it
that could serve a vTPM.
The other arrangement puts a small trusted component at VMPL0 and the guest kernel at VMPL1, and that component serves the vTPM from inside the encrypted guest. Azure does this with its paravisor, and COCONUT-SVSM is the open version. There the host cannot read the TPM’s state, at the price of putting that component inside your TCB, where attestation had better be measuring it. Same device and same code, but in the first setup the platform can read the key and in the second it can’t. Find out which one you’re on before you describe this property to anyone.
wrapping up
None of this stops an attacker on the machine from using the key; while they are there, they can sign whatever they like. What they can’t do is take it with them, so their use of it ends when their access does.
A private key in a file is a bearer credential that happens to sit on a machine. A key in the TPM is closer to a property of the machine itself. It’s also the only answer to post-attestation machine identity I’ve found that doesn’t undo what the attestation just established.
What’s missing is the link between the key and the attestation, and for that I’m looking at SPIFFE. Roughly: the vTPM’s public key goes up alongside the evidence, and a short-lived certificate over that same key comes back down, an X.509-SVID. The private key never moves. The certificate expires and gets reissued while the key underneath it stays put, which is the same split as earlier in this post.
The mTLS at the far end of that is the code you’ve already read here, pointed at a different CA. I’m writing it up next.
Working notes. If something here is wrong, tell me.