Security Cryptography Whatever
Security Cryptography Whatever
The feeling's mutual: mTLS with Colm MacCárthaigh
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
We recorded this months ago, and now it's finally up!
Colm MacCárthaigh joined us to chat about all things TLS, S2N, MTLS, SSH, fuzzing, formal verification, implementing state machines, and of course, DNSSEC.
Transcript:
https://securitycryptographywhatever.com/2021/12/29/the-feeling-s-mutual-mtls-with-colm-maccarthaigh/
Find us at:
https://twitter.com/scwpod
https://twitter.com/durumcrustulum
https://twitter.com/tqbf
https://twitter.com/davidcadrian
"Security Cryptography Whatever" is hosted by Deirdre Connolly (@durumcrustulum), Thomas Ptacek (@tqbf), and David Adrian (@dadrian)
So what I'm hearing is it's a zombie, so you can stick a fork in it, but aim for the head, so it's dead permanently.
SPEAKER_00Hello. Welcome to Security Cryptography Whatever. I'm Deer Girl.
SPEAKER_05I'm David. I'm Thomas. And I'm Columb.
SPEAKER_00Yay! Tom's our special guest today. I'm a cryptographic engineer at the Zcash Foundation.
SPEAKER_04I am an engineer at a company called Name Tag, but I also did a PhD in almost cryptography at Michigan and co-founded Census.
unknownCool.
SPEAKER_05I'm an engineer at Fly.io and I have one semester of undergrad college.
SPEAKER_03I'm an engineer at Amazon Web Services where I work on uh security, cryptography, identity, and virtualization. This is very exciting.
SPEAKER_04You're one of the coveted principal engineers at Amazon, correct?
SPEAKER_03Yes. I'm definitely a member of the Amazon uh principal engineer community. It's fun. Welcome, Column. What what is that? So at Amazon, you know, when you join as a as an engineer, typically you come in as as uh an ST1, right? Which is uh typically a college hire. And then ST2 is our next level at that. After that, it's kind of a still kind of an early career position where where people learn how to be really great software developers. Then ST3s are lead developers, and then after that, you become this you know blessed principal engineer who you know supposedly we we are great at you know stewarding and shepherding projects across teams and across the company and figuring out what the technical direction should be as well as getting real things done. Does it come with powers? Uh it comes with a title, it doesn't, there's no intrinsic powers. Uh it comes with responsibilities. I I found out the hard way that you know tends to make things harder rather than easier.
SPEAKER_00But do you have minions?
SPEAKER_03I do not have minions. No. It's I I work with work with a lot of teams and I have to help a lot of engineers, but none of them just do my bidding. Everybody always wants to argue and figure out what the right thing to do is.
SPEAKER_04How much bar raising do you do per day?
SPEAKER_03Uh all the time. Everything. It's kind of it's a constant bar raising, it's more of a life philosophy than anything else, you know.
SPEAKER_05I I gave a talk at Amazon like a bunch of years ago about like, I think it was our cryptography for pen testers presentation that we had done, like at a bunch of places. And when I was there, like it was like a blue hat kind of thing where they had um outside people from a bunch of places just presenting to the Amazon team. And like a bunch of people at Amazon, like the engineers there were talking about how they would get like sponsorships or recommendations from people outside the company. Like that was a strange Amazon culture thing I was unaware of. At least at the time, there was some process where like it mattered if somebody outside of Amazon like wrote you a recommendation or said something nice about you. Do you have any recollection of that? Am I crazy?
SPEAKER_03You're not crazy. I it's it's definitely something that comes up. We do, you know, when people are doing annual review processes or promotion processes and all that kind of stuff. It's not that unusual to look for feedback from folks who don't work at Amazon, in part because you know, we try to be really customer obsessed and getting feedback from customers is super valuable, but also partners and vendors and other people we work with because you know their feedback's really useful.
SPEAKER_00This reminds me of the tenure process when you're going up for tenure and if you're in academia, you have to get letters from people outside your department or you know, outside your school who could be like, Yeah, I would totally give them tenure or whatever, but I'm somewhere else. So yeah.
SPEAKER_03Yeah, and and I think when when part of someone's role is, you know, interacting with the industry, it becomes particularly important. I can definitely think of some people like that's most of their job. Yeah. So they probably have all sorts of feedback and recommendations like that.
SPEAKER_05I just want to tell people working at Amazon that I am available for recommendations and that my fees are very reasonable.
SPEAKER_00Do we want to talk about S2N, like the AWS cryptographic library?
SPEAKER_03Sure. Amazon S2N, it's it's an open source project. It's on sound GitHub. You can find it there at our AWS repo. It's short for signal to noise, which I I'm amazed was a name that was left there in terms of for you know for for a cryptographic library, how no one had taken that name before us, you know, because the core function of cryptography is to turn meaningful signals into useless noise, right? And to to hide information in plain sight. But it's it started as a library that just implemented TLS, right? And SSL. And we would use, you know, OpenSSLs, lib crypto or other lib cryptos under the hood to do the core cryptography. But steadily, bit by bit, we're kind of taking on more and and even implementing that. We've we've got a our own lib crypto project as well that that kind of plugs into it, and we've grown to support QUIC and some other stuff in there. We've got some uh cryptographic primitives that we use in non-TLS contexts as well, that's all part of S2N. And I think it's been going six or seven years now. I can't remember.
SPEAKER_05I'm trying to remember what was going on with TLS six or seven years ago. Like what was the impetus for what was the impetus for starting your own TLS implementation?
SPEAKER_00Something other than OpenSSL.
SPEAKER_03So develop- I mean development on S2N started literally the day after Heartbleed. Just to get directly at that. So that that was definitely a trigger. But we had talked about and discussed our own internal need for S2M before that. And we had actually outlined it and we were planning on starting it about probably about six months later than we did. Because we saw there were like some performance optimizations that were kind of on the table. And we saw that we just had to get into this game of owning this part of the stack ourselves for a bunch of reasons. But, you know, when Heartbleed happened, it definitely accelerated my timeline. I literally, literally started working on it pretty much full-time right right after that. Not quite S2N itself first. The first thing we wrote was actually um a kernel module that would act as a network filter that could block Heartleed. Because we had some customers stranded who couldn't update their copies of OpenSSL. So we wrote a little module they could use on their instances. And then that kind of grew to become S2N.
SPEAKER_00So they couldn't upgrade OpenSSL, but they could install a kernel module?
SPEAKER_03Yeah. So we we had a bunch of customers who were kind of in unusual situations. Some we were able to help them with hot patching, right? So heart bleed was a pretty easy issue to hot patch binaries for. You know, literally just find a block of code near that processes these heartbeat requests and add a jump and then add a section at the end of the jump that, you know, adds a condition to defend against heartbleed and then jump back to where you went, right? It's actually not that hard to do. But we we had some customers who had, you know, validated binaries that they couldn't change. They had self-checks and checksum that have to pass. And so you couldn't modify the applications themselves. Yeah. So you so you can't hot patch it. So now you got to do something in the network. And that meant you know, shadowing and parsing the entire kind of SSL TLS state machine, detecting your heartbeat record and rejecting it. So that's what we did.
SPEAKER_00Is that code still live deployed?
SPEAKER_03Um, I hope not. Uh the module is still public, it's in my GitHub repository, but I hope nobody's still running that. I that was the that was definitely you know intended as a bad day to to help some people get by who who didn't really have any other option.
SPEAKER_05I'm looking at like the introduction for S2N, the like your announcement post, how it played on Hacker News in 2015, because that's my lens for how to look at everything. Yes. And the top comment there is about a 12,000 line OCaml implementation of TLS. And my question for you is why did you not implement it in OCaml?
SPEAKER_03Well, I've I personally am not OCAML literate. So I implemented it in C. And I guess the the the other big reason to implement it in C at the time is pretty much everywhere we did TLS, you know, every application that did the the front-end SSL TLS processing, it was also written in C.
SPEAKER_01Yeah.
SPEAKER_03And so we didn't want to be constrained. And we had some, you know, compilation target environments that couldn't even didn't even run, couldn't even compile C to. So we we were pretty restricted. We were able to get to C99. Pretty much all even our embedded environments could support that. But that was pretty much our lowest common denominator at the time. You know, it wasn't what I do today, but that's that's the that's what we did at the time.
SPEAKER_00What kind of embedded environments are these clients talking to AWS or something else?
SPEAKER_03So as the name is, it's Amazon S2N. And it's not just used in AWS. Yeah. And in fact, at the time when we wrote it, one of one of the some of the smallest compilation targets included things like dash buttons, which are you know, literally I don't know if you can remember those. Right? So think about trying to get an it, you know, a TLS deck that can run an environment like that. Got it. Very, very small, very tiny footprint.
SPEAKER_00Oh, oh, so there's like a little section of post quantum crypto and the s2n refile. Are you deploying post quantum crypto to dash buttons? Please say yes. Please say yes.
SPEAKER_03Wait, no, I don't think so. As fun as that would be. But uh, I I don't know if you can still buy them.
SPEAKER_01Huh.
SPEAKER_05This is like the thing where if you're out of detergent, you just have a button next to your washing machine, you push it, and then detergent comes.
SPEAKER_00Yes. Yeah, they're really cool. They are really cool.
SPEAKER_04I I actually got one that just every time I hit it, I get a new TLS implementation.
SPEAKER_03Um yeah. It's uh you know, now now you can just ask your echo device to do it for you, and you don't even have to press the button. So it's a good button.
SPEAKER_05So like Heartbleed happens. I so I guess you announced S to N in in 2015, but Heartbleed was like years before that, right? But like you have a sense of where like OpenSSL is now, but you also have a sense of where it was back then, right? Like it's easy to sell me an S to N in 2012 or whatever, like in the in the bad old days of OpenSSL, right? But like if you were gonna make a sales pitch right now for like use this different C implementation of TLS instead of OpenSSL, uh, I believe that there's a good pitch there. I'm just wondering what it is.
SPEAKER_03Well, first, I don't mean to criticize OpenSSL and don't ever don't ever take anything I'm saying. I think OpenSSL is probably one of the the greatest world goods that is ever achieved in software development. Literally, like a mostly volunteer team that brought cryptography to the masses. It's just an amazing accomplishment, and I don't ever want to talk that down. I've also pretty good friends with a lot of the OpenSSL team, so I don't want to get in trouble with them. We all we all agree, which which frees us to be mean to it.
unknownOkay.
SPEAKER_04We can put in the standard disclaimer now, too, though, like OpenSSL in 2014 and OpenSSL like now are also miles apart. Yeah. I also wouldn't, you know, in OpenSSL 2014, I wouldn't go like beat anyone with a stick because of heartbleed. I think there's a a lot of things that kind of kind of led to that happening. Yeah. But it's certainly what you're talking about when you're creating ST1 then is is not at all what OpenSSL is like now.
SPEAKER_03Yeah, good good context setting. So so the big motivation for us was really to have fewer security issues to deal with. And there's there's two senses of that, right? One sense is, well, we we really do have to have a lower risk profile. There are, you know, sometimes bad actors coming after our customers, and we we got to be able to protect them, and there's some real threats there. But the other level of it is no matter what, anytime an issue comes out, anytime, no matter how low the risk is, you got to go do a bunch of updates. And those can be pretty disruptive. At Amazon scale, rolling out a software update in a low-level library or a low-level system, you know, like SSL and TLS is used everywhere, like can cause, you know, a lot of teams to have to pause. And then they, you know, don't work on their own roadmap for a bit. They go and pivot and they have to do this update and get it deployed and do all their testing and so on and so forth. And I don't know a good way to measure the impact of that, but it is certainly tens of millions of dollars at Amazon scale, right? It's like just a huge uh over many years, the amount of whatever that productivity loss is. And the development cost of something like S20 is always going to be less than that. And as long as we can do it in a way where we add the right defenses in depth and have a much more minimal surface area, the advantage is you get got to just sit there and write out all those updates. And that's pretty much worked. A lot of up issues have come and we just didn't have to do anything at all. And all of that productivity is saved. And that's the the main benefit, and that was the main pitch, and that's how we kind of continue to go at it. Then secondly, the performance. You know, there's performance enhancements that are still on the table. You know, we still have ideas for how we think we can do pretty serious performance savings and how this stuff is uh, you know, works. And we we got some pretty big wins, you know. When you're looking at a system like Amazon S3, even like two, three percent performance improvements turn into you know big numbers very fast.
SPEAKER_00Yeah, especially when you're like the the first stop for a cache for anything for the internet to stick it in S3. So any of that when you're effectively a cache for the internet, any of that latency is felt, including in your SSL handshakes or you know, whatever.
SPEAKER_05I guess I I mean I I should have asked like a second ago, but like where does S to N live in the in the architecture right now? Like literally everywhere in AWS where there would be a TLS, is it now S to N?
SPEAKER_03Not quite everywhere. I mean uh there's there's a few left. We're very, very careful about how we update and making sure we don't break customers and and we preserve backwards compatibility along the way. But pretty much every if you talk to an Amazon Web Service, if you call it API, that's S to N. If you hit the CloudFront CDN, that's S to N. If you hit Amazon S3, that's S to N. If you hit network low balancers, that's S2N. If you hit an application low balancer, it might not be S2N. That's one of the few things that's left on our list, but we're working on that one. And that'll be fun.
SPEAKER_04In terms of the productivity kind of gains, are that coming because you have to patch less or because you can patch easier because it's like within your build system and owned by you?
SPEAKER_03Yeah, I think the answer is yes. And um it's it's mostly having to do nothing. You know, it's it's nice to be able to say, well, this issue came in, but we don't have to worry about it. We just don't have to do anything. That that is by far the biggest win. And then because it is in our development and build system and it's it's kind of a first-class project internally, that is easier. It's also, you know, if there is an issue with S to N, right? If there is a security issue with S2N, we're generally, you know, gonna be where any security researcher reports it to, right? So we're gonna have kind of first hand privileges in the embargo process, right? And be able to coordinate with them and have everything updated and and the the day customers find out about it, everything's already batched, you know? So it's a a good productivity win too.
SPEAKER_00That's an interesting interesting way to be like, is there a value to a pro a a quote unquote fully open source project like OpenSSL that's available to quote anyone? But that means that it's also controlled by no one in particular. So everyone has to coordinate and you're at the whim of the project and you have to try and get your fixes in line. And if you're an Amazon, you're just like, well, no, that I can't I can't come, I can't work in that workflow. I'm just gonna go over here and do it on my own. So is there I guess the value would be to smaller organizations than an Amazon or an AWS or Google.
SPEAKER_03For for us, we we maintain a full Linux distribution too, Amazon Linux, and and so we we already have to patch everything, you know, well within embargo timelines. And we we have to be able to do that. And so and we we put a lot of work into being able to respond to any issues that are reported in any kind of project, yeah, you know, very, very, very quickly. That is a hard thing for a small organization, I can't even imagine.
SPEAKER_00Yeah.
SPEAKER_05I guess like I feel like I know the answer already, but I'll ask anyways, right? Like if I was to come at you as like the skeptic saying, I simply don't trust memory unsafe C software, like what is the set of things that you guys have done? And I know there's a bunch of things you guys have done to mitigate that concern.
SPEAKER_03So, well, firstly, I I'd I'd point out that almost every memory safe language, especially the dynamic ones that you can think of is itself written at C at some level. You know, we have exceptions now that are able to be self-hosting, but you know, at the time, even if you you thought about the JVM or or huge things like that. And so there's no like magic that they have access to that you can't also do in a C program. So we essentially wrote in our own dialect of C, like a pretty restricted dialect of C that takes its inspiration from functional programming techniques. And so we we structure all of our memory handling and I. You know, called stuffers, where we create these things, you know, which is like a buffer for stuff, right? And it it's really, really simple. It's a very simple data structure that keeps a cursor, right? And anytime you write to it increments the cursor, and anytime you read from it, make sure you don't read past the cursor. But when you use a technique like that and you write code like that, it makes everything beautiful and declarative looking and very functional seeming, right? We did another example of that in how we constructed our state machine where we're literally we're using like a fixed, you know, table of function pointers, and all you do is increment your way through the table. And so it's kind of writing C like you'd write lift. So maybe maybe your own question was it wasn't so crazy. And then we did, you know, just an enormous degree of testing and verification all the way up to formal validation.
SPEAKER_00Oh, I've heard about that, yeah. So it sounds like the answer is you have to write your own variant of C to do this securely.
SPEAKER_03I don't know, you have to, but it's what we did. I always think of you know, your starting point is the first layer of defense is always gonna be writing the code and whoever's reviewing the code. And you have to make that as mentally untaxing as possible. And you have to make things as consistent and idiomatic as possible so that anything unusual will really stand out. And for me, that means you probably want to go for a pretty restricted small set of patterns that you're gonna program with and use and start there. Yeah, you know, and then make tests easy to write and have lots and lots and lots of testing and test cases, you know, and then do fuzzing because fuzzing is really, really awesome. And then after the after you've done all those things, you know, if you can spare to, you know, think about doing some formal validation as well. What is the formal validation story there? What are you guys doing there? We formally validate quite a bit of S to when we we do things like we we formally validated uh our state machine that it can't get into any invalid states that aren't allowed by the the TLS specification. We formally validated our memory safety by doing some some processes. We have a tool called CVMC that we run on uh like our core stuffer uh and IO algorithms and all that stuff, and we we validate that that's always going to be correct no matter what the input is. Nice. We formally validated our implementation of HMAC.
SPEAKER_00Cool, yeah, I remember this. Yeah, yeah, yeah.
SPEAKER_03Yeah, that was fun because we we wanted to see if we were getting better at formally validating things, because about two years prior, there'd been another formal proof of a different HMAC implementation. We wanted to see if we could make it easier. And we formally formally validated a bunch of our algorithms, more and and most of the post-quantum ones that it that we're working on as well.
SPEAKER_00You have formally validated them? Like bike and psych?
SPEAKER_03We are formally validating those.
SPEAKER_00Cool.
SPEAKER_03I think those algorithms aren't yet like locked in to the point you can even say that they're they're validated, right? They're still open to tweaks on various parameters.
SPEAKER_01A little bit.
SPEAKER_04You're validating that the implementation matches the algorithm or that some property about the algorithm itself.
SPEAKER_03So most of our validation, what we do is we validate that the actual compiled like machine code matches some simple specification, like declarative specification of a protocol or a safety property, you know, some kind of invari set of invariants that we we want to be able to hold it to. But we we try to go all the way to the machine code where we can.
SPEAKER_05Not it. I'm I'm gonna nerd out a little bit here, right? But I'm I'm just curious about like what the tooling looked like and what those projects looked like.
SPEAKER_03Yeah. So Ga Galois was one of the companies we partnered with super early. Actually, not long after we started S2N, we had um Byron Cook join AWS, who's now a distinguished scientist at at Amazon. Cool. Awesome guy, real leader in the field on formal verification. And he connected us with Galois so we could get going before he was able to hire a full team. And we still work with GAWA, they're awesome. They have these tools called Cryptol and SAW that are for, they're specifically designed for verifying cryptographic code and cryptographic algorithms, which can be, you know, some traditional verification tools from like the safety critical world don't really apply because there's so much entropy and randomness in cryptography, and you have to be able to kind of abstract that out. They were able to come up with, you know, cryptol and saw specifications for the HMAC algorithm and for the core TLS state machine. And they they were able to find issues. It was really cool. They actually found that there were certain like invalid combinations of TLS extensions that could cause us to abort the TLS state machine early. Now it wasn't a security issue, but it was still uh pretty pretty good find. There's I I can't think of any other way we would have found that.
SPEAKER_00So is that the the specification allows this? Can be conformant with TLS like one three or whatever, and you have all these extensions, and that is allowed by the spec. But if you implemented the spec as written, it would get into this weird state that you really should not be allowed to be in.
SPEAKER_03So TLS supports session resumption in its state machine, right? So normally when you connect to something over SSL, there's a handshake, there's a bit of a back and forth, and you eventually negotiate a key and then you encrypt stuff over it, right?
SPEAKER_01Yeah.
SPEAKER_03But you can skip most of the handshake by doing resumption. Yes. And what they found was if you showed up with this weird combination of extensions in which you shouldn't resume, it could try to resume, and then the resumption would fail. And what it should do is just carry on and do a full normal handshake. But it it instead tried to resume, didn't resume, and kind of gave up and aborted the handshake early when it shouldn't. That was that was the issue they found, which would be really hard to find.
SPEAKER_00That sounds like a a section needs to be added under session resumption in the TLS 103 spec to be like, well, make sure that if you're, you know, you know, blah blah blah blah blah blah blah. So did anything get updated after you found that?
SPEAKER_03Yeah, so so that was on the state machine that's part of, you know, SSL V3, TLS 1011112. The TLS1.3 handshake is radically simplified.
SPEAKER_00Okay.
SPEAKER_03All right, and does not suffer from from any issues like that. There's just nowhere near as much kind of parametrization of the the TLS handshake in 1.3. It's way, way simpler. And and that was inspired by issues like that that you know that Gawa found.
SPEAKER_00That's awesome.
SPEAKER_04And the the 1-3 handshake or specification itself has been symbolically formally verified. Think so. The spec doesn't have invalid states with Tamar.
SPEAKER_00Yeah, yeah.
SPEAKER_05I guess having been through the process of taking a relatively complicated state machine like TLS and then seeing it formally verified. And you're like a you're a software developer like the rest of us, right? Like, are you still comfortable building state machines, building protocols, like just building software without formal verification? Like, I have friends that do this and they never come out from behind the looking glass, they're just permanently like formal verification people. And from that point on, they just don't take seriously anything that isn't formally verified. Have you fully drunk the Kool-Aid on that now?
SPEAKER_03Um, may maybe not fully. So what I'd say is first, if you have to build a state machine, right? And if the first rule is in your code, in your functions, do not mix input parsing and changing states, right? Don't like mix those core things, like separate those really clearly. Try to have like a set of code that really clearly describes how you're gonna go through your states and have that very separate from the code that's gonna like parse input and do stuff like that. You know, a lot of projects, they're a mix of those things. You know, there's these huge big functions that do both, and it gets just too confusing and gets too spaghetti-like and is unreviewable, yeah, in my opinion. So you first gotta do that. That's like the biggest thing, and that structure helps. The second thing then I tell people to do is linearize all your possible state transitions, right? So instead of having like, you know, conditions in your code that can go, well, if this happens, then jump to this state and so on and so forth. Instead, lay it all out in these linearized tables of like, well, this is one valid set of state transitions in like one table. This is another valid set of transitions in another table. If you have too many tables, if that feels like you won't be able to program that, give up and redesign the state machine. Like that's a that's a big hit that you know you're you're too flexible in all your state transitions. And those two things I think matter more and probably prevented more issues and bugs for us. And then formal validation after that kind of gets you to 100%. It's like you do those things, you'll get to like 80, 90%. But you know, formal validation is probably the only way you're gonna get to 100%.
SPEAKER_00This gives me a little more confidence about doing Zcash stuff with the network and the this and that in Rust. And we basically we are doing as you described, we have parsing over here and it throws errors. If it doesn't do anything, it doesn't understand. We have state transitions and how we handle them over here, but also shout out to Rust. There's the the type safety in Rust makes it very easy to encode states as Edom variants or types, and to say you can only go from one state into another valid state. You cannot go from any to any or you know into to from valid to invalid or something like that. And you can check it at compile time, and it's very nice for that sort of thing.
SPEAKER_03So big fan too. You can you can use S to N with Rust, right? Like there's bindings. Yeah, there are bindings. You can use S2N for most time, we do. We have various Rust projects that that use S2N, including the new Rust STK for for AidWest. And we go the other way. There are now parts of S2N that we're writing in Rust. Like our quick implementation is is written in Rust because we feel like it's ready and it's a better starting point for all that.
SPEAKER_00That's so exciting.
SPEAKER_05What are the prospects for Rust fully infecting the S to N project and you're gradually hoisting out most of the C code?
SPEAKER_03Um not imminent. In gener in general, we try to leave code that's working alone and not go rattle it. But I mean I'd love to see it.
SPEAKER_00In what, 2014?
SPEAKER_05Well, I mean they could say it now and it's much more credible because like lots of stuff has happened.
SPEAKER_03But I mean, it'd be great to see it someday, but we don't have any imminent plans.
SPEAKER_00I meant to ask, does AWS have a fuzzing cluster, or are you leveraging open open fuzz or whatever it's called?
SPEAKER_03So on Amazon, those kind of practices are kind of up to each team and what they want to do. But we do we do have some centralized fuzzing infrastructure and we do have a compute cloud, it's called EC2.
SPEAKER_00Yes, you might have heard of it.
SPEAKER_03It's got some compute, it's got a it's got a few instances we can use now and then. Cool. And um we certainly do use it. I mean, we we've been fuzzing on S2N for it's been running for years and years and years at this point.
SPEAKER_00Awesome. Did you write your own management to run and report and correlate, or are you deploying? Because I tried to deploy whatever they there's a pro cluster fuzz, OSS fuzz, the Google run one is they have a piece of software that's like you know, Kubernetes. Here is how you have a web app that deploys your your fuzzing infrastructure, and then you can tie it back to your, you know, your repository or whatever it is. Do you have something like that?
SPEAKER_03Well, the first fuzzing tool that I wrote, I just use Elastic MapReduce. Okay. Um and kind of like trick through it at it that way. I've seen us use Lambda for it as well. Really? Just as a nice cool demo. I think Lambda's cool for fuzzing. Obviously, you can't run things for very long, but it's still useful for integrating it directly into the to build process if you want to get just a really, really quick yes now on something. Um and we we do that sometimes because it's not feasible to run it on everybody's desktop for that kind of stuff. But uh I don't know actually. I should ask the team what they're how they're coordinating and running it these days. There's a lot of different ways to run something in parallel across many EC2 instances.
SPEAKER_00Yeah, but like specifically, like a lot of people will write a fuzzer and then they're just like run the fuzzer for, I don't know, some amount of time as part of their CI or CD or something like that. And then you're not really getting the real benefit of fuzzing. You need like a continuously running fuzzing infrastructure, and then you also need it to report when it does something, and then you have to correlate it back with the change that actually you found the thing on. And all of that work, not just the like deploying lots of compute in parallel, turns out to be only one solved problem, at least as I've seen. Actually, that's not true. There's cluster fuzz if you run it yourself. Someone took cluster fuzz and they were trying to run it as like a service, not they basically OSS open source fuzz, but like pay them to do it well for you. And I think they shut down or something like that, and it made me very sad. So if you have software that makes this task easier for people to do, or you know the people who have it, I am interested and I would like to see more of it in the world, please.
SPEAKER_03Sure. I'll I'll ask them. Maybe maybe we should have a service for it. I always view the the stuff that's integrated right into your CI pipeline is is really just to give the developer feedback that they haven't broken the fuzzing infrastructure.
SPEAKER_00Yeah. But you want to auto-deploy anything you've changed with to like your fuzzing infra infrastructure or whatever. So either way. Enough about S2N. Tell us about MTLS.
SPEAKER_05Oh wow. Actually, I'll put something in the middle there, right? Which is just like you've now had the experience of being like firsthand to a you know, ground-up implementation of TLS. I think I bring this up every time the subject comes up, right? But like one of my favorite people is Watson Ladd, and Watson Ladd had like a comment that has stuck with me forever on the CFRG mailing list, which is the IETF crypto review board, where he compared TLS to like an undergraduate secure transport, like undergraduate homework assignment, and said that you would have gotten a C on it if you had turned it in. What's your general take on TLS at this point?
SPEAKER_03Well, I think it takes all sorts of internet standards are like that, right? And you can't be too hard on them because you know, a lot of them came out through a culture of experimentation and iteration, right? And then sometimes something takes off and succeeds wildly before maybe the you know people got another chance to iterate, and then you're all stuck with it because you know it's just baked into everything.
SPEAKER_02You know, too.
SPEAKER_03But one of my favorite examples of that is actually like TCP and UDP, like even like way back, like UDP has this crazy design error where it does fragmentation at the IP layer, right? Like the header in a UDP packet is only in the first packet of a fragmented datagram, which complicates you would not believe the amount of extra money that makes routers and switches cost, right? Because they have to be able to like reassemble those packets and so on to be able to do flow switching, right? And you would look at that and you go, well, you'd get a C minus on that design. It would have been trivial to just put the UDP header in in every packet, right? But you can't see it like that. It solved the problem and it did it really well and took off. And TLS is the same. I mean, you can look at it and say, Oh my god, look, they got the Diffie helmet exchange the runway around. I mean, this is like this is this is clearly a C minus C minus. But you know, they built a really cool protocol that could effectively emulate TCP enough that you could just bolt on existing protocols and get going, you know, and it's it's a it's a good example of an FVP succeeding. And maybe we got better, we have to get better at iterating, and it shouldn't take 20 years before we're able to like all come back and you know, let's hammer out a better, more optimal design. So I kind of see it like that.
SPEAKER_05I feel like there's a point where you warned me about the UDP fragmentation thing. We did like an all BPF implementation of UDP for fly, and like you were like UDP fragmentation, you didn't use these words because you don't use words like this, but you were like UDP fragmentation is gonna screw you with like DNS sec, and I'm like, Oh, it's gonna screw up DNS sec. Oh no, and then I moved on. That's only because you mentioned that only because you mentioned it now do I remember that you warned me about that. So, like part of the reason we're bringing MTLS up is there's like a sort of brave new world of how people say microservices, and I hate that term, but like modern application service ensembles. That's MOS MACE, that's my new acronym, anyways. For these MACE applications, right? Like, there's this notion that like you've got all these services running now, delivering the same application. They all talk to each other, right? And like we now have an opportunity to use TLS to secure the connections between those services, right? Like you've got all these random things talking to each other, and like in the bad old days, there'd be no good way to kind of authorize who's allowed to talk to what. It's all just kind of like you'd run TCP dump to see what's going on. And now what we can do instead is like bolt a proxy onto everything and have it talk MTLS and like MTLS is the way that you would, like TLS, the TLS protocol that we're talking about, right? But in mutual mode where you're presenting both client and server certificates for things, right? If you look at that, like you can get a long ways into the kind of the authorization and authentication problem, you can get a long way solving those problems just by using certificates and both sides of TLS, right? And that's essentially it's kind of where Kubernetes is going, right? Is towards something that looks like that. And I gather you're a great fan of this.
SPEAKER_03I I am not. And I think when I when I first saw the designs from from like Istio and Spiffy and so on, I sent them a very long note with like my 56-point detailed critique of why you really don't want to use MTLS for this. But you know, at the same time, they're solving a problem, right? And they're plugging into a layer that they can. But I'll I'll I'll try to give some more detail. I have a long history with MTLS. When I was still in college, one of the ways I was kind of paying the bills, I was a member of the Apache HTTP project. Cool. Writing code for for the Apache web server. And at that time, it was still mostly non-US folks who were working on the SSL stuff because of you know silly crypto export restrictions and so on. And so I would help people with the SSL stuff, and I would help developers with them, and I would help, you know, people who are just running Apache with them. And I would I would do these workshops and and through that got into this kind of you know business of being an amateur auditor of people's MTLS setups. You know, they would they would come to me or or or through somebody else at Apache and say, can you take a look at this and see if we've actually done it securely? And first I was not a professional security auditor or a viewer or anything even resembling that. So the fact that they were people like coming to me and I was pretty close to maybe one of the best people out there to do that was a really bad sign that this was not a very mature, you know, ecosystem. And uh literally in every single case I looked at would find unbelievably low-hanging issues, like like stuff that it just didn't work at all. And on top of that, it was you know, it's a really complex ecosystem. When you're when you're using MTLS, there's a lot of X509 flying around, a lot of strength comparison flying around. And like I was saying earlier, like you're gonna have to respond and update to a lot of security issues when you bake that really deep into your stack.
SPEAKER_02Right.
SPEAKER_03And that's that's gonna slow you down. But like just some of the top things. Literally, a revocation almost never had they built a working revocation system. And I kind of think about, you know, when people tell me you should back up your data, I'm like, the first question I'm asking is, well, but how do I restore it? Like that's the important part, right? And when somebody tells me you got to be able to rotate your passwords, I'm like, I don't really care about rotating them. I want to know how you revoke them. Like, and and tell me, you know, how do you make sure something can't be used again? And they would just almost everyone would like, well, we'll build that later. Or, you know, they would use CRLs or some huge list, and then I'd ask them, well, what if you have to revoke everything? Like, is that gonna scale? And they just never really have an answer, you know? And it was kind of scary. And then I I would find these cases where authentication wasn't happening at all. You know, people couldn't tell. Like they were using the system, they were going to some internet sites in their in their web browser, and you know, everything worked, but under the vote, the client certificates weren't even being used, they were just getting regular TLS or there's no easy way to tell, you know. And I found cases where people were doing authentication just based on the strings that are in X509. My favorite example of that was one of the first projects I did. We found that the the CTO's EA, right, his executive assistant, could pretty much do anything she wanted. Like she had God level power in the system. And at the time I thought it was, well, we we must have put her in the CTO's group, and the group CTO's group has has root power, and and so that makes sense. We'll track it down and figure it out. But it wasn't quite that. It was literally just that she had the word admin in her job title. Oh, and they were just match pattern matching strings and regex that was those looking for admin and and didn't didn't didn't have the dollar sign terminator and stuff like that. And it's it's just full of that. Now, these things that are building on it now, you know, these these mesh networks and so on, you know, they're they're much more professional and they've they've taught about a lot of that and they're they're compensating for a lot of that, but they still mostly still have X509, you know, stuck in there pretty deep. And so you're gonna have to update for every X509 parsing issue that comes out, you know.
SPEAKER_05I feel like revocation is kind of where you sold me on this. Like you you wrote a long thread on Twitter about your MTLS gripes, which I am going to shamelessly plagiarize in a blog post at some point, right? But like we had written like a blog post kind of cataloging different inner service authorization things, and we said fond things about MTLS, and you said unfond things about MTLS on Twitter, right? And like I think going into it, I might have been prepared to put up a fight. And then you talk about the revocation thing, and it immediately clicks for me that like the revocation stuff in TLS that we're familiar with and that we talk about is not the same problem as like inner service revocation or even kind of any kind of like API revocation, right? Like they're just different problems, right? Like when we talk about like internet scale revocation, we're talking about generally targeted attacks or specific mississuances and things. And like there's there's a sense in which that system kind of converges on correctness over time and you kind of hope things shake, you know, shake themselves out. And anyways, any attacker that's going after that system, like they're a passive adversary that has control over traffic, anyways, and all that. But like in API authentication, you have to be able to revoke, like people not revoking, right? Like you lost a credential. Like, if you can't revoke it, then people can keep like forever using that credentials. Like the system permanently loses security. And I feel like I look at like how just how these systems are built, right? And I get the sense that they think that they're drafting off of a lot of security that TLS has that TLS never promised to provide, right? Like there was never a notion that that that TLS was gonna solve, you know, fine-grained, immediate, real-time revocation the way that we expect, you know, even like OAuth or something to do.
SPEAKER_03Yeah, the revocation has always been the stickiest problem in in TLS on both sides for for client certificates and for server certificates. And there's some some genuinely hard problems in there, and and maybe there are better ways to solve it outside the TLS kind of ecosystem. For our inter-API or inter-service auth that we decide for AWS, we decide to do everything at request level, right? We authenticate every specific request. So in TLS, you're authenticating the channel, right? And then you're just blessing the channel, and anything that happens over that channel inherits the auth. And that means you can't authorize and authenticate specific transactions at the same time, which you know, we just feel is a very weak security model. So we we just go for it at the request level. And then we use, I mean, we have uncountably large sets of identities and credentials. You know, we we issue very ephemeral identities and very ephemeral, you know, session credentials that can last seconds, hours, you know, and so you could never even try to do that, something like that with TLS. So some you can kind of try by baking it in epochs, you know, some you can kind of myt short-lived credentials and say, well, this expires at a certain time or past a certain epoch, but that doesn't give you the ability to revoke on demand. And also if you're able to isolate something or if you're able to influence time, and often things are just using NTP for time, you can overcome that too. So it's it's full of all these little gotchas. So we just kind of went, no, do everything at the request level, and we're just gonna use, you know, HMAC and symmetric keys and go at it that way.
SPEAKER_04So let's say though, like just to play the other side a little bit, like if you are authenticating on every request and you have revocation, that means that like more or less you're doing like a revocation check on each request. And so is there any reason that you couldn't just do OCSP on every TLS connection? Like you're you're you're paying the check every connection cost either way, so why not just pay it in TLS?
SPEAKER_03Yeah, so we we don't quite do a revocation check on every request. Instead, we kind of do proactive full lifecycle management of every credential, right? So when you create a session credential, you push it out there, right? And it gets to the places that can authenticate. And when you want to validate or evoke it, you do the same thing, right? And there are some fallback safety measures. If something becomes isolated, it knows it's isolated and stops serving requests and so on. But in general, it's a live, you know, positively acknowledged feedback system, which is really, really important, right? Because if you're making changes, right, like you don't want to use a new identity or credential until you're sure everything that could authenticate with it has it, right? And in the opposite direction, you don't want to stop using one until you're sure it's no longer in use, right? You don't want to kill it unless you're like, you know, sometimes there are cases like let's say one of our customers has fired an employee in in negative circumstances. In that case, they do want to, you know, break their access very, very quickly. And you can do that. But if the system's pushing things in general, it's not having to go do checks. So you get efficiency and you get the management too.
SPEAKER_05I feel like there's stuff worth thinking about in the channel versus message thing as well, right? Like probably one of the most important kind of server side attack vectors right now, or I'm I'm like a year out of date because I've been doing just pure software engineering for the last year. But like when I stopped doing assessment, like one of the major things that was like, you know, important for server side, like when we were doing assessments and stuff, was things like SSRF, things where you have like, you know, an existing channel and then you know, just being able to send a message. Over it, it's counterintuitive, right? Because like you wouldn't imagine that just being able to turn a server application into a proxy would be that big of a deal. Like you can get a proxy anywhere, right? But like if you've got a trusted HTTPS channel that it's already trusting and that anything that goes over it is is blessed, right? Then you've kind of got game over anytime anybody gets a way to slip a message into that channel. And you don't have that problem with like SIG v4 message authentication or stuff like that.
SPEAKER_03Correct. And I think probably D Sync is probably even the newer kind of uh form of that, right? You can you can imagine a D sync issue if it occurs at one of these proxy layers that's using MTLS to bless the channel, will will have that issue in a way where an authenticated request will not.
SPEAKER_05I'm guessing that like James Kettle hasn't tested that yet. Because no one's really testing Envoy and stuff for like, you know, you can't write a scanner or you can't make a like make a BERT pro you know, a BERT plugin that doesn't I'm saying this and somebody's gonna point out that it exists, right? That's a smart thought, right? Like there's probably is something there.
SPEAKER_03Yeah, it's it that's probably potentially a target-rich environment, but we can't do it from that perspective too, right? Like there's all sorts of ways you can end up breaking a HTTP request and putting strings in places they shouldn't be, and you know, bad escaping and so on. And so let's be more defensive there. Well, the one we were thinking about is well, we want to be able to target things like, you know, banking apps and and financial customers and so on, where they literally want to sign the specific transaction and sometimes even want to sign it offline because they're they're very paranoid with their keys. And it's very hard to do that with MTLS.
SPEAKER_04Right. So let's say you sold me on why you shouldn't use MTLS for like microservice authentication. But what about in like a kind of employee or device authentication or like zero trust scenario? But let's say SSH to make things simple. Like, let's say I have servers, then I have employees. And while you can certainly make the argument that nobody should have SSH permission in anything, let's say you need that type of permission, and you're like, okay, what I'm gonna do is I'm gonna build like a key vending machine that's gonna give out client that's gonna use my organization's SSO to give out SSH client certs, and then people will authenticate to whatever server they want to using this client cert. And that way my config management only needs to push out, you know, the verification cert for SSH. This also kind of gives me a little bit of an edge in the argument since uh SSH doesn't use X509 certs. Yes. And like you can make me you don't need to explain to me why X509 makes things complicated. But like, what about a scenario like that that is an API authentication? Is are client certs just fundamentally broken, or can they be used in other scenarios?
SPEAKER_03I don't think client certificates are fundamentally broken. I think PKI, you know, is is useful and SSH is an example where it is useful. I've definitely seen customers build setups like that, although maybe they're not as common as they should be. Most people are still just using, you know, either SSH passwords or or just generating a public key and dropping it on their box. But um, I think a lot of my arguments against MTLS fall away in that case because they're now, you know, no longer you're no longer just trying to create this like TCP compatible, you know, pipe or tunnel that almost anything can run over, and you're just gonna ignore, you know, the context of that protocol. You know, SSH is very coupled and they've really thought through how very carefully, you know, the implications of public and private keys and how that impacts the protocol and how that affects things and so on. And I think there's all sorts of other great uses of PKI systems for building, you know, identity frameworks and giving people, you know, long-lived identities. Sometimes it's it's really the only way to do it, like with physical cards and so on, if if you want to be able to work with systems that work like that, and they definitely have their place.
SPEAKER_05I guess like I might stick up for it in, you know, I might stick up for MTLS itself in a couple of scenarios, right? Like um, we use some of the HashiCorp stack at Fly, right? So there's the there's console in the mix and there's some nomad in the mix for us, right? And um, HashiCorp is really big on MTLS authentication. So I think Go programs in general are MTLS positive because Go makes it pretty straightforward to use client certificates, right? And in those settings where I'm basically expressing what is sort of kind of a network topology to begin with, like the relationships there change with my network topology. Like there isn't a whole bunch of like issuance going on and stuff like that, right? And a thing I kind of like about it as opposed to like fine-grained request authentication is that if you don't have the client certificate, like if you don't have the root secret for talking to the service or whatever, you can't talk to it at all. You're just like kind of locked out, right? And that that's a thing that MTLS does that I do kind of like, right? It's similar to similar to SSH, I guess, too. If you've turned password authentication off, then almost everything that people write about SSH hardening goes out the window as well, right? Like unless you don't trust the SSH key exchange and stuff, which hasn't been broken in forever, right? Like I do sort of like the idea of you know really coarse-grained kind of you know access control rules expressed with MTLS. You could tell me I'm crazy about that. You should, by the way, if I'm crazy about that, tell me that I'm crazy about that.
SPEAKER_03Well, I I've definitely seen it fail open more times than I've seen it fail closed, where people, you know, just stand up a web server, think they have mutual auth working, and it turns out they don't. But I, you know, the Go ecosystem is much better at that. It's way more explicit. I can, you know, it's it's harder to get wrong than than an Apache config or or an Nginx config where it's really, really pretty intimidating recipe to make sure this thing's actually turned on.
SPEAKER_01Yeah.
SPEAKER_03So that part I agree with. I mean, at that level, it's kind of logically identical to a VBN, you know, and sometimes VBNs make all the sense in the world too.
SPEAKER_00It's dovetails.
SPEAKER_05So we were talking before we started recording, and we were talking about like roughly what we were to talk about. And you slipped in right before we started recording that if we put you on the spot, you might try to make an argument that AES CBC was more secure than AES GCM, right? So CBC is like old school AES. It's how like AES protocols were designed, it's how block encryption stuff was done since like the 1990s, right? Where in particular you've kind of you separate out the encryption part and the authentication part, you kind of compose them from primitives kind of generically, right? And GCM is like, you know, it's it's like a Formula One car, right? The whole thing is just hermetically sealed around doing both authentication and the the bulk encryption at the same time, right? And like I think a lot of people like GCM. I'm not a big fan of GCM, it seems really brittle to me, but I I am a fan of AEAD ciphers, like you know, like make the argument, convince me that I should use CBC.
SPEAKER_03Oh my god, give me two things to argue against there, too, because I'll have to get back at you about A D as well. So TLS kind of famously got AES C B C the wrong way around, right? There's you can't use AES, C B C on its own. You have to couple it with some kind of authentication algorithm or a Mac, right? And TLS uses HMAC, but you know, if you give it a plain text, it would do the HMAC first and then encrypt the HMAC, which is the wrong way around. That's not what you want to do. It means somebody can, you know, now means the ciphertext is malleable, somebody can mess with things and try to do experiments. And we got security issues out of this, most famously the looky13 security issue, right? Which Kelly Patterson and co-found, which is awesome, and showed us, but definitely don't do it that way around. Now you can use it the other way around securely, right? If there was a cipher suite defined that, you know, when you're decrypting, check the HMAC first and then decrypts it. You know, I I don't think anybody out there would actually would would have any standing issues with with ASCBC, except except maybe performance. It's not quite as performant as ASGCM. But ASCBC has like a lot of positive properties that if you're thinking from the perspective of like a real-world attacker, somebody who might really try to go at your protocol, it's it's really defensive. You know, the the biggest one is if you screw up in how you generate your initialization vectors or nonces, which are just these things it takes to uh to encrypt with it, it's much, much more defensive of that. In fact, GCM can break wide open in a way where you'll be able to forge information. And secondly, it gave us a measure of length hiding that I think was really important. And so the cryptographic research community out there, you know, their job is to advance the art and think of, you know, the next cool break in cryptographic protocols. It's not to defend real world systems from real world attacks, right? And it's staggering to me that the most practical attack on TLS to this day is well, if somebody's passively tapping traffic, whether that's you know, shared Wi-Fi or whatever, they can see the length of all the information that's going by and the rate of the information that's gone by, right? Just classic traffic analysis stuff. And any measure of length hiding helps protect that. And there's like real practical attacks here, you know. People have done research, they can they can see, you know, what map you're looking at in the browser because they they figure out the mapping tiles or what video you're watching on a video streaming service because they figure out the size of the movie segments. Or, you know, if you're pushing VoIP around, you can just watch for the the silences and the breaks and figure out what speech patterns are, right? And most ordinary people would be really surprised to learn that. And be also be really surprised to learn that like these encryption things don't actually protect their information in that way. And at the same time, most cryptographers would be surprised to learn that they are surprised. You know, they'll be like, of course you could do that.
SPEAKER_05Just to make sure that I'm I'm tracking this. I think I am, right? You're getting the length padding from CBC because CBC is padded. Because when you put together a ciphertext with CBC, it's gotta be a multiple of 16 bytes long. And if it's not, it's you fill out the balance with garbage.
SPEAKER_03Exactly. And it's a relatively minuscule amount of padding. Like I'd prefer a much larger amount of padding, and now TLS 1.3 has support for larger amounts of padding, but no one's using it yet. But it's pretty even that 16 bytes, it's actually really effective at hiding the URLs, right? Like when you when you think about trying to analyze a HTTP session, you've got two attempts, right? You've got the size of the request and the size of the response that you can use to try to fingerprint what's going on. And it really makes the attacker job much, much harder on the request side because a lot of URLs will collapse into that same 16 bytes. And and this stuff has, you know, it increases attacker difficulty by by many, many, many factors, you know. And it's just a shame to lose that, you know, over an HMAC. Yeah. It's like we're not thinking through the perspective of like, but sit down and try to do a practical attack on TLS, what would you do? It'd be something like this. And it's like, well, we actually went backwards a bit on defending against that, which is just kind of perverse. It's just just a little weird.
SPEAKER_00So my instinct there is that the block cipher mode or an authenticated encryption with associated data mode is like a different level of abstraction than HTTP over TLS length extension or you know, length observability attacks, blah blah blah. Like if you are designing your block cipher mode or your AEAD thinking about that, like are you kind of like are you out of your element? Are you gonna try and shove too many sort of things into your thing that you're designing at a far lower level of abstraction of the primitive or quote unquote primitive that you're designing? Or basically should you consider that because I don't know if CBC mode was ever I doubt it was designed with that sort of security when integrated into a higher level protocol in mind, it just was a happy accident, basically.
SPEAKER_03Yeah, it was definitely accidental security. So I definitely think padding should be part of the AAD kind of API. And and myself and Shy Guran designed an AAD spec two years ago called Scram about for AES. And we well, padding right there, front and center. So be in the developer's face, you need to think about how much padding you should have, right? And here's why that matters. And you see it ill-considered, you know, in in a fair amount of applications and protocols out there where folks really haven't thought thought about, you know, just basic blinding information to attackers. So I think I think we have some more work to do there. But you can't put absolutely everything a developer has to think about in the AAD, you know, fing fingerprint either.
SPEAKER_05Hey David, you've got your name on some TLS papers. How convinced are you?
SPEAKER_04I mean, I've been generally skeptical of most like timing type attacks. I don't know about traffic analysis as much, but most anything that falls under the side channels as to like what is revealed there that simply like existence of the connection doesn't reveal. There's spots where it's like clearly terrible, right? Like the old Dawn song at Berkeley and SSH with the whole keystrokes on the password thing from back in the day, right? Right, that's clearly an issue. I'd believe that like many audio formats might leak this type of stuff. And a lot of those are SSH, for example, has like specific things in the implementation to avoid the keystroke timings. And I'd believe that like even if you have the 16 bytes of padding and audio, you might still need to do things like that. So I don't know. I was never much of a timing person. I agree that you know anything you can do to mask lengths and mask stuff is positive. You see this with the censorship resistance stuff a lot more, that just like anything that can be used as a side channel and that does get used as a side channel, and sometimes like really silly stuff, or even just the attitude of like, well, surely they'd never block AWS and then they like block all of AWS type stuff.
SPEAKER_05So I don't know. You're bothering me. How seriously do I need to take the SSH keystroke timing thing? Oh, that's been that was fixed like 20 years ago. Well, right, right. But like it's implement, it's it's implementation fixed, right? Like, but if I'm using some random, you know, SSH library in a memory safe language where like I I trust the crypto, but like there's a whole bunch of other things you have to think about when you implement SSH. Is that one of them? Like, do I need to look for that?
SPEAKER_04That's probably the first one you should look for. I I say this because like I'm tangentially involved. I'm involved on the um on the transport crypto side, not the SSH side of like a kind of SSH re-implementation for reasons. But I I don't know, I would believe that the I haven't checked this in the Go one, but I assume that the Go implementation does this correctly because Go has a lot of smart people that worked on that. And like I just I personally wouldn't touch drop pair with like a 10-foot pole, regardless of the keystroke timings. And that's about all of the SSH implementations are SSH drop pair, and there's maybe one other one. And it basically, I mean, I wouldn't touch anything that's not open SSH or the Go one, is my answer to you. Literally, as you say this, I'm remembering that I wrote our own SSH implementation for Fly, and then you can go look for this and but it it it imports the Golang slash X one, which is like the same thing that Okay, yeah, but it's a it's an it's an X library, they're not the same.
SPEAKER_05I'm gonna we need to bring Philippo on like right now, right now. Yeah.
SPEAKER_03Summon Filippo. You gotta watch out for the the Kly utilities though, that's the problem, right? The the way it works in SSH is the second you go into line buffering mode, right, or password input mode where you get the asteriskes, it tells your terminal over SSH to like not send the input until the the cards return, right? But if somebody writes an application that doesn't go into line buffered mode but accepts a secret, like it's it's game over and you still see those.
SPEAKER_05You can distinguish it. I've seen that in ostensible secure messengers before where they send real-time updates like who's talking, like somebody is typing, but like every time they type, like you get like fine-grained messages like when the keystroke actually happens. And I always wonder like how crazy that is. Like, you know, you probably want to do something to like debounce or jitter that or something like that.
SPEAKER_04Yeah, I mean, I saw stuff back in like 2014 of like figuring out what you typed just by listening to it.
SPEAKER_00So like I hope that would be good things.
SPEAKER_04I'm sure it's not great.
SPEAKER_00Coming back to CBC or other ADs, I think that if you came up with another AEAD and was like, hey, it's all padded so that you can't distinguish length, someone would be like, it should be secure without all that extra padding, you know, blah blah blah. Why is it so fat? And it would be harder for you to like get support for your thing if it was thickly padded, even if it was to mitigate exactly this kind of attack when it's used in a higher level protocol.
SPEAKER_04I mean, he saw like exactly this with like AE80s the first time around. We can all we can go and look up that email from Phil Ragaway from like 1996, yes, where he's like, You motherfuckers are fucking all of this up. He would never say that he's like one of the most timid people, but that's what he's saying. Basically describes everything that goes wrong in the next 20 years, right? Then and then they're like, eh, but performance, or like, and this isn't realistic. So he's just saying, like, use that AEAD abstraction, which like had just he had basically just come up with it.
SPEAKER_00And then like 20 years later, everyone comes back. We're like, We're sorry, can you like release the patent on this? We really like this.
SPEAKER_04Yeah, or same with like, you know, let's build custom block, like we have custom block cipher mode specifically for disk encryption, because that's a separate problem. It doesn't seem unreasonable to me to want to have a separate block cipher mode or AEADs that are designed for length hiding to be used in transport layer encryption.
SPEAKER_05Okay, um, like like that doesn't like that seems reasonable to me. I know the thread that you're talking about, right? That's the one where like people were referring to Phil Rogoway as like a so-called cryptographer when he was trying to fix IPsec. All I have to say about this is that obviously the next time we do a recording, one of the things that we need to do is a dramatic reading of that email with you as the anger translator. I'm not kidding in the least, it'll work. Yeah, it'll be great. I have just like one more question for you, Colin, which is in a day that shall live in infamy on this podcast, we had uh Ryan Sleavey, who is one of my personal heroes, and uh I uh kind of with some hubris brought up DNS Sec and asked if it was okay for us to stick a fork in it and declare it dead. And uh he gave a whole speech that has shaken me to my core for the last you know two weeks about how what they're doing in Europe with certificate authorities means that there is a future for DNS Sec. And in fact, you know, we may be soon in a Dane world. And I would like, I'm not even asking you a question so much as making a request, which is would you say he's wrong, please?
SPEAKER_04And it's important to note that that request was not individually authenticated.
SPEAKER_03We authenticated as transport connection first, so I think DNSEC I it's probably best described as a zombie right now. In in it is it is the living dead. And I say that with a twinge of regret, you know. I think you know at this point I feel like DNSEC is mostly being propped up by regulation. At least the folks who I see using it tend to be the folks who have government mandates to use it or ICANN mandates to use it, which I guess is is just another form of government. And it isn't getting too much commercial adoption. And I'm kind of skeptical of the protocol's fundamental future just because it seems really poorly positioned for like the next round of updates or break, you know, it's really, really hard to update the parameters in DNSEC, very, very slow. Some of that's due to the nature of the protocol and just how crufty all the things out there that support DNS are. If if you think the the challenges that you know Greece had to overcome in the TLS space are anything, they're 10 times harder in in the DNS world where there's all sorts of very old DNS implementations out there that really can't handle any changes.
SPEAKER_01Yeah.
SPEAKER_03And you know, things will come along and updates will have to happen. And what are folks going to do? And then the other part of it is there's just there's not much of a you know deep technical community around it who are able to steward and shepherd all that. You know, TLS has a really deep bench of of folks who work on that, who, you know, got to work every day and that's their job. And and they keep the world moving on that stuff. And DNSEC doesn't really have that. So it's really hard to see it kind of thriving. You know, at the same time, I don't want anyone listening to think, you know, we don't stand over the DNS sec or or MTLS invitations we have at AWS. We support these things and we put a lot of work into getting it right for people and and making sure that they they work really, really reliably. But you know, these these bother ecosystem concerns are real too.
SPEAKER_05I feel like you hear sometimes, like, you know, a thing that gets brought up with DNSSEC a lot is that the ecosystem as deployed right now is mostly RSA 1024. There were like recommendations to use 1024 for like the mainstream keys, and then the longer-lived key signing keys can be like, you know, 2K RSA keys. But like the obvious criticism is just it shouldn't be RSA at all, it should be curves. And then people will point out, well, there is curve, you know, DNSSEC, right? Like you can use the p-curves of DNSSEC, and Cloudflare has an implementation of p-curve DNSSEC. And it's like there's like two problems there, right? You know, the obvious problem there is that p curve DNSSEC is like, you know, 1990s curves, but like the other problem is just like none of the deployed, you know, the actual deployed base of TLS or of DNSSEC is p curve DNSSEC, right? Like you imagine how long it would take to get that stuff deployed, and you know, it seems like it'd be an eternity to get that attack to actually happen, let alone getting like curve 25519 or something like that deployed.
SPEAKER_03Yeah, and it's been at least 10 years just to get 1024 out there, and that's gone pretty, pretty slowly. I don't even know where you would start on if you wanted to really deploy EC. And you'd have to turn off OSA as well, really, to get to get the skill. benefit. Like it's not, it's really two steps. And getting to the end of those two steps, I mean, at the current track record would would take over 20 years. So that's why I'm skeptical of it. But you know, it's a protocol, it's got security in the name. Folks are gonna want that extra comfort of using it too sometimes. And you know, at least we're getting better as an industry of of avoiding outages and, you know, operating it very well. We're starting to see more mature implementation show up that, you know, safeguard against that. So there's less downside to using it than there was before.
SPEAKER_05When Ryan was making his case for this, I think I should have been more forceful. And I think you're seeing me attempting to be more forceful about my dislike for DNS sec right now. So what I'm hearing is it's a zombie. So you can stick a fork in it but aim for the head. So it's dead permanently.
SPEAKER_03Well I I personally would love to see a revived effort to have a real you know secure DNS protocol. I actually think there's if you step back from the internet architecture and you were doing everything afresh, right, you would want encryption to be like a day one property, right? You'd want to you'd want to be right in there and you'd want it to be part of the name lookup, right? You you kind of want everything that happens in DNS and everything that happens in the TLS hashtag to be one protocol. Right. And you just you start with a name and you get back an IP address and a key right and maybe a port whatever. And then you go connect to it and you're good, right? And doing that well with you know confidentiality and privacy which DNS Ec doesn't have right you know with full verification and and authentication at each stage it would be really really awesome. Like there's definitely a space for that and hopefully we'll see it at some point.
SPEAKER_05You're starting to see it bottom up with things like DOH, right? But like if you step back and look at like how this came to be right what you're really just seeing is like sheer bloody minded path dependence, right? Like we're working with the constraints that a small dood project had in like the mid-1990s where like it's not encrypted because they felt like DNS servers of the time wouldn't be able to keep up with encryption right and there's a there's a notion of offline versus online signing where like the protocol can't you know make any nods to online signers like you know anything that actually has a key and can do cryptography in real time because I mean you can make a message board argument that that's a good property but the reason it's there is because they felt like the systems of the time that were gonna you know when DNS was fully deployed in 1997 which were whatever the plan was right like those servers wouldn't be able to keep up with it and we need a protocol yeah that's designed around 386 SX you know DNS servers or something like that. Right. I totally like I buy completely like I come across as an evangelist for not securing the DNS but I think DNSSEC gets in the way of securing DNS.
SPEAKER_03I agree with that and I think I'm optimistic that we can solve all those technological problems and challenges. I think there's another property too though which is maybe speaks to your other concerns about DNSec and PKI about you know they can be subverted by single parties sometimes right whether that's a a a DNS operator somewhere in the tree or or a rogue CA or whatever right and I think stepping back you'd you'd kind of look at it you go well wouldn't you love to just make that a multi-party system where you know everything has to be signed by end parties and that's computationally cheap now and and we can have many CAs now and so on. But that kind of change where essentially everyone becomes dispensable right by design those are harder to do. So the kind of changes you're ambitious for I'm I'm I'm less optimistic about but we'll see. So you're saying it's a coin no I don't think blockchain is the answer.
SPEAKER_05That's like you said secure multi-party computation and I'm like you have dear does attention so yeah do I have to write another threshold signing vertical for DNS I'm not even talking about threshold signatures here.
SPEAKER_03I'm I'm just talking about you know you could you could have to say you know something in the DNS or a certificate right just has to be signed by multiple parties right and then you could just have the software exactly simple simple as that right would would be another interesting change.
SPEAKER_00Well Colin thank you so much for coming on our little show this has been great. Anything else?
SPEAKER_05No it's been great too cool awesome it's great to it was awesome to finally meet you and talk to you. Thank you so much for taking the time thank you