about things notes.zzstoatzz.io
notes
notes sources talks.md
51 kB

the way that it does. 0:02 2 seconds Sweet. Can everybody hear me? Okay, good stuff. Um, cool. I'm going to be talking about appro ethos, which is 0:09 9 seconds essentially the philosophical and aesthetic principles underlying the design of the protocol. So, before we get started, my name is Daniel Homegrren. I'm the head 0:18 18 seconds of protocol here at Blue Sky. Uh, you can find me on the atmosphere at dhomes.xyz. And for the freaks out 0:26 26 seconds there, you can find me at this did So the title for this talk uh actually 0:33 33 seconds came from a thought that I have had whenever I saw Martin post this. He said whether something is decentralized or not is a function of the administrative 0:41 41 seconds control of different parts of the system not a function of the network topology. 0:45 45 seconds And this just struck me so hard and I was like yes this is exactly what we're going for. I actually brought this up to Martin a couple of days ago. He didn't even remember posting it. I'm like I think about this every week Martin. 0:58 58 seconds uh but I quoted this and I said this is appro ethos and uh I do think that this is very descriptive of the topology and 1:06 1 minute, 6 seconds the power dynamics in the network uh but what it doesn't quite have is like a prescription for how to get there and so this term ad protoethos kept coming up 1:14 1 minute, 14 seconds in our protocol design questions we kept kind of like nice uh we kind of kept coming back to it whenever we were weighing different 1:22 1 minute, 22 seconds options and I wanted to take the chance to try to dissect what that ethos is the historical influences on the design of the protocol, some of the new things 1:31 1 minute, 31 seconds that we brought to it, as well as some as well as some of the postures that we adopt as we are working on the design of it. Uh, and just for what it's worth, 1:39 1 minute, 39 seconds not all of this was like perfectly packaged or anything from the beginning. 1:42 1 minute, 42 seconds We had like a million false paths and dead ends and hard debates between team members and everything. And maybe in a future talk, we can give the philosophical history of app proto, but 1:51 1 minute, 51 seconds I'm going to give you guys the nice clean packaged version of it. 1:55 1 minute, 55 seconds So I would sort of situate ad proto at the center of uh three movements or three trends and influenced by those three trends. On the one hand we have 2:03 2 minutes, 3 seconds the web. Uh then we have the peer-to-peer movement and we have I don't have a good quick word for this but kind of large distributed systems or 2:11 2 minutes, 11 seconds these data inensive applications that kind of run the modern web. I also didn't have a good picture for that so I used the bore from Martin's book. Uh so 2:19 2 minutes, 19 seconds we'll start with the most obvious influence which is the web. Uh, as I'm sure you guys all know, the web is an information system for publishing documents over the internet. So, if you 2:26 2 minutes, 26 seconds have a static IP address and you have a server, you can host a document. If you have an internet connection and a route to that server, you can fetch that document. Uh, kind of the incredible 2:35 2 minutes, 35 seconds thing about the web, the defining feature of it is that it's completely permissionless. That's the only thing that you need to publish a document. 2:40 2 minutes, 40 seconds That's the only thing that you need to go fetch a document. Location is also based in the authority. We we layer on DNS on top of those IP addresses. This 2:48 2 minutes, 48 seconds sounds really obvious to say out loud, but you know that a document is the thing that you get from a server because you talked to that server and the server gave you that document, right? But the 2:57 2 minutes, 57 seconds authority rests in the location that you talk to whenever you're fetching the document. Uh this combined with the permission list of the web was actually 3:04 3 minutes, 4 seconds a pretty radical idea. You didn't need any central publisher, any central distributor. You didn't need a certification checklist to say that you 3:11 3 minutes, 11 seconds could publish a document or anything like that. If that permissionless and openness was the defining governance feature of the web, sort of the defining 3:20 3 minutes, 20 seconds product feature of the web was the hyperlink which gave you interconnected data. Uh documents could link to other documents and specifically they could 3:28 3 minutes, 28 seconds link across authority again completely permissionlessly. The downside of all of this is that servers are hard. These all run on servers. Not everybody knows how 3:36 3 minutes, 36 seconds to run a server. Not everyone can get a static IP. Not everybody can configure DNS. And so what we started to see emerge in the modern web or what people 3:43 3 minutes, 43 seconds call web 2.0 is this move from permissionless publishing to platforms. 3:48 3 minutes, 48 seconds Everything running on a big server called bigocial.com which I thought about making a BS joke but I thought that might be dangerous with the name of our 3:57 3 minutes, 57 seconds company. Uh so these these platforms were account based. You'd register an account with one of these platforms. The benefits of it is that it's very easy to 4:05 4 minutes, 5 seconds post. You don't need any technical knowledge. You just need an account on the platform. you have rich media, uh, photos, videos, music, interactions, 4:13 4 minutes, 13 seconds dynamic real-time interactions between accounts on that data. And then to top it all off, a nice algorithm that churns through all of that data, uh, shows you 4:21 4 minutes, 21 seconds exactly what you want to see, crunches all those like counts, the reply threads, everything like that. Uh, and I don't want anyone to think that I'm like 4:29 4 minutes, 29 seconds doy nostalgic for web 1.0. There's a lot about the modern web that is really great. Nobody wants to go away from that, but there was also something that 4:36 4 minutes, 36 seconds was lost with it. So that kind of takes me to the next influence on at proto which is the peer-to-peer movement. In a lot of ways this was a response to what 4:44 4 minutes, 44 seconds was happening in the modern web and that sort of centralization in these large social platforms. And the basic question that it asked is why is there a server 4:53 4 minutes, 53 seconds in the mix? So I have me here and Devon over there. I'm writing a post. Devon follows me. I want to send that post to Devon. And we say why is the server 5:00 5 minutes here? What if we just removed it and I sent the post directly to Devon from my phone to his phone? And the question that came out of that is can you build a modern social network without a backend? 5:10 5 minutes, 10 seconds Uh this kind of came from a distrust of these large centralized uh server services and platforms. And so if I if 5:17 5 minutes, 17 seconds uh Paul, Devon, and Brian are all following me, I make a post. I send that post out to the three of them directly from my device to their device. Of 5:25 5 minutes, 25 seconds course, there's a problem there. If I'm not online at the same time as them, how does my post actually reach out to them? 5:31 5 minutes, 31 seconds And so these peer-to-p peer pro uh protocols said,"Well, if I sent my post to Paul already and I'm offline and he's online at the same time as Devon and 5:38 5 minutes, 38 seconds Brian, then he can forward on the post to them." Of course, the problem with that is that what if Paul replaces my post with a poop emoji and sends that to 5:46 5 minutes, 46 seconds Devon, and that's really embarrassing for me. I don't want Devon to think that I posted that. And so these protocols would get around that by introducing public key cryptography signing the 5:56 5 minutes, 56 seconds message. Uh which essentially put the certification that the message is what I intended it to be in the message. Every 6:03 6 minutes, 3 seconds account would have a public key associated with it and it gave me the liberty to broadcast my message out and trust that other people could deliver it reliably with all without altering the 6:11 6 minutes, 11 seconds contents of it. Uh then in a lot of these protocols you would then get content address data. 6:18 6 minutes, 18 seconds uh you would sort of build this abstraction around the routing of these messages. Since you could get them from anywhere, the certification is in the data itself. You would reach out and you 6:27 6 minutes, 27 seconds would say, I know that there's this piece of data with this hash and you would reach out to the network to go fetch that and grab it from some device in the network. What this gave you 6:35 6 minutes, 35 seconds whenever it was functioning well is this really beautiful abstraction. It's what drew me to the peer-to-peer movement in the first place, which is it completely dissolves away the network topology, the 6:44 6 minutes, 44 seconds devices, the locations, everything like that. And there's like kind of this essential essence to these posts that are kind and you can just reach into the 6:52 6 minutes, 52 seconds ether and pull them out. It's like a really beautiful thing. It's a really beautiful concept whenever it's working well. Of course, that's if the abstraction is working. A lot of the times it doesn't work well and you have 7:01 7 minutes, 1 second uh data availability problems, data routing problems, uh things like that. And then the big problem is that phones are a lot 7:09 7 minutes, 9 seconds smaller than data centers. And we have these big machines in the cloud that are crunching all of this data for us. And phones just cannot keep up with that same data. you have resource usage 7:17 7 minutes, 17 seconds problems, battery life problems, and then you just can't store all of that data on your phone. So, you're not going to be able to create the same global views that these big social platforms 7:25 7 minutes, 25 seconds are able to create. Um, excuse me. 7:32 7 minutes, 32 seconds So that brings us to the last influence on approto which is the parallel development that was going on while these peer-to-p peer protocols were gaining or were kind of being um worked 7:41 7 minutes, 41 seconds on which is these how these large distributed systems that underlied the modern web were being architected. So 7:50 7 minutes, 50 seconds sort of the features of these is that they have millions and eventually billions of users uh real-time dynamic views. Someone likes a post you get an update on that milliseconds after the 7:58 7 minutes, 58 seconds fact. They're very very readheavy. For every post that is posted, it might get loaded thousands or millions of times. 8:04 8 minutes, 4 seconds And there's very high expectations whenever it comes to downtime and latency, especially as more and more of the interaction on the web started centralizing on these 8:12 8 minutes, 12 seconds platforms. So, we needed new techniques for building these data inensive applications. Uh Martin Kleman's book designing data intensive applications is 8:20 8 minutes, 20 seconds kind of the bible on this. He gave a really great overview of all of the trends. Uh I'm going to talk about kind of high level the trends. Obviously, all 8:28 8 minutes, 28 seconds of these platforms are architected uh differently, but maybe we can sus out some of the commonalities in a lot of them. So, maybe one of these platforms 8:35 8 minutes, 35 seconds starts out with a server and a big Postgress database. Uh, and the first thing that you're going to run into because they're so right read heavy is 8:43 8 minutes, 43 seconds that you want to split out the read and write loads. Uh, so here we see that database getting some read replicas. All of the writes are still coming in on a 8:51 8 minutes, 51 seconds leader database. The reads are getting pushed out to these read replicas. you can scale up reading separately from the rights that are coming 8:58 8 minutes, 58 seconds in. A lot of these platforms started to decompose their monolithic services into microservices. You would essentially peel off some piece of the service uh 9:07 9 minutes, 7 seconds wrap it in its own service. The monolith or the main service would then call into it. There are a lot of reasons to do this. Organizational reasons. You could 9:15 9 minutes, 15 seconds have small teams own code bases. Uh you could also scale resource usage according to what that service did. 9:22 9 minutes, 22 seconds video transcoding looks very different from timeline generation and you could put these on machines and give them resources that matched what they needed 9:28 9 minutes, 28 seconds to do. Um, of course, whenever we do this, all of a sudden we just created a distributed system. And whenever you 9:36 9 minutes, 36 seconds have a distributed system, you have distributed problems. Uh, you no longer have strong consistency necessarily between all of these services. If one of 9:44 9 minutes, 44 seconds these microservices goes down for three hours and it comes back up, did it just miss out on three hours of writes? Is it ever going to get back in contact with 9:51 9 minutes, 51 seconds the same data that the other services in the data center have? And so we find these services starting to lean into eventual consistency uh and oftentimes 10:00 10 minutes eventually or oftent times uh using eventually consistent databases as well. 10:05 10 minutes, 5 seconds I put Sila here because that's what we use at Blue Sky. Uh but it's a successor to Cassandra. There's other data databases of that type as well. And 10:13 10 minutes, 13 seconds these kind of give up on the acid guarantees of Postgress or something like that. give up on the ability to uh linearize writes and uh really focus on 10:22 10 minutes, 22 seconds being able to do low latency writes and reads. Another aspect of these is that you also give up on the ability to dynamically query your data. You kind of 10:30 10 minutes, 30 seconds have to be aware of what your application is going to ask of the data uh before you lay it out in your database. That means that if you need a new index or you need to reindex stuff, you have to reprocess all of the data. 10:41 10 minutes, 41 seconds If the data is already in the database, how are you going to do that? A lot of these problems were generally solved by an approach of stream processing. I 10:49 10 minutes, 49 seconds called out Kafka here, but again there's other examples of uh that. And this kind of introduced a information flow into these backends where writes would come 10:56 10 minutes, 56 seconds in, they would get added to uh the stream processor and then filter out to the microservices or the database. This meant that if a microser went down for 11:05 11 minutes, 5 seconds three hours, it could pick back up, talk to this canonical stream of data and catch back up, know that it has everything. Similarly, if you have to 11:12 11 minutes, 12 seconds rebuild an index or build a new index, you can just replay the entire stream of canonical data and know that you have everything in that index. These kind of 11:20 11 minutes, 20 seconds these canonical stream processors sort of served as the backbone to these data centers uh that could uh give the 11:27 11 minutes, 27 seconds feeling of consistency in them. So, this brings me back to approto ethos. Uh I would say that at proto ethos is sort of 11:35 11 minutes, 35 seconds situated as the synthesis of these three movements. Uh so it's influenced by all of these movements. From the web, we kind of took the fact that it's open, 11:43 11 minutes, 43 seconds permissionless, and it's this universal uh network of interconnected data. From the peer-to-peer movement, uh we took 11:51 11 minutes, 51 seconds location independence of data, the fact that the data is self-certifying. It has the certification embedded in it and this general skepticism of 11:59 11 minutes, 59 seconds centralization. And then from these large data inensive distributed systems, the splitting of read and write load uh 12:06 12 minutes, 6 seconds large application aware indices decomposition of stuff into microservices owning up to eventual consistency and uh stream 12:14 12 minutes, 14 seconds processing. So now I'm going to get into the uh new things that appro brought to the table. The first is identity based authority. Again, this is in contrast to 12:23 12 minutes, 23 seconds the location-based authority of the web or the contentbased authority of peerto-peer. We said identity is in privacy. So the central thing is my 12:31 12 minutes, 31 seconds account at dhomes.xyz and then the thing that you'll find there is all of the social data contained inside of it. With this 12:39 12 minutes, 39 seconds came a change in addressing scheme. So in the web we have https. It points to a location bigo.com and then you fetch my 12:46 12 minutes, 46 seconds profile from that. A peer-to-peer protocol like ipfs points to a uh hash of some data. In app proto it points to 12:54 12 minutes, 54 seconds dhomes.xyz XYZ which kind of looks like a location but it's actually an identity and then it points to the uh record type the collection and the key of the record 13:03 13 minutes, 3 seconds inside of that. Similar to peer-to-peer protocols this gives you location independent data. Uh every identity even if it looks 13:10 13 minutes, 10 seconds like a DNS name resolves to a universal identifier that is independent of location and then that identifier in 13:18 13 minutes, 18 seconds turn resolves to the location where the data is stored and key material. So we essentially uh introduce this layer of indirection between the identity and the 13:26 13 minutes, 26 seconds location of the data and that lets these identities kind of float on top of the data abstracted away from the actual architect or the actual topology of the 13:35 13 minutes, 35 seconds network. Uh because they're floating on top of the network they're able to migrate between locations that actually host the data or host that signing key. 13:44 13 minutes, 44 seconds Um and uh identity is really the primary thing that floats at top this layer of hosts. 13:51 13 minutes, 51 seconds And then from the web we kept interconnected data but and similar to how the web has interconnected data across authorities at proto also has 13:59 13 minutes, 59 seconds interconnected data across authorities but those authorities are identities and the data is structured uh schematic data instead of document 14:08 14 minutes, 8 seconds data. So the next thing that we introduced is this notion of generic hosting and separating application semantics out from canonical data 14:17 14 minutes, 17 seconds hosting. So this is my identity that we saw earlier. It has some likes, follows, posts, and images inside of it. From the 14:24 14 minutes, 24 seconds perspective of my data hosting server, my PDS, it just looks like this. They're just Seabore objects. It has no idea what the rich application semantics 14:32 14 minutes, 32 seconds inside of my repository are. It just goes, this is this is Seabore. What this lets happen is that application developers can then own 14:40 14 minutes, 40 seconds their schematic space without coordinating with the host of that canonical data. So inside my repository 14:47 14 minutes, 47 seconds we have uh blue sky posts and smoke signal events and whitew blog entries that are defined by the application 14:56 14 minutes, 56 seconds app.bvents.smokes signal com.whitewwind not by the host that is hosting that canonical data. All of these uh hosts of data then 15:05 15 minutes, 5 seconds exist in this very crawable, very accessible network of open data and applications can then reach into that data, pull it out and imbue it with the 15:14 15 minutes, 14 seconds rich application semantics that you associate with the modern web. So what this gives us is decentralized hosting and centralized 15:22 15 minutes, 22 seconds product development. The hosts that make up this decentralized network are blissfully unaware of all of the rich applications that are running on top of 15:30 15 minutes, 30 seconds it. But the rich applications can make choices that let them move like how you expect products uh to move. They can have uh you kind of need to be able to 15:39 15 minutes, 39 seconds have leadership whenever you're making a product. You need to be able to rapidly iterate. You need to be able to use the tools that you want to use, but you inherit all of the benefits of the 15:47 15 minutes, 47 seconds decentralized hosting and the decentralized data where it can be extensible, remixable, reusable, accessible to other applications while 15:55 15 minutes, 55 seconds retaining the benefits of having a uh a uh clear leader for a given product. So the next two things that I 16:04 16 minutes, 4 seconds want to go over are less like features or something that we introduced and more kind of postures that we took as we were approaching these problems. Uh the first 16:11 16 minutes, 11 seconds is kind of the acknowledgement that having structure actually gives you freedom in one of these open and peri per permissionless networks. So approto 16:19 16 minutes, 19 seconds is a multi-party lowcoordination network. You have all of these devices and services talking to one another and the question is how do they all talk to 16:27 16 minutes, 27 seconds each other and how do you come up to a coherent hole. You want to avoid the tyranny of structurelessness which is essentially this collapse whenever you say hey it's really great we have 16:36 16 minutes, 36 seconds freedom. we can do whatever we want and you have this collapse in coordination where nobody can actually get anything done because there's no clear leader to actually get that done. Uh we want to 16:44 16 minutes, 44 seconds spend our time building applications and doing cool stuff and not just focusing on facilitating interoperation, fixing weird edge cases, building defenses 16:52 16 minutes, 52 seconds against bad actors or trying to coordinate evolution whenever none of us have a uh clear leader that can do that coordination. So what you'll find is 17:00 17 minutes that at proto is a very structured protocol. We tried to put uh we tried to reduce the number of choices that you could make in the important areas in as 17:07 17 minutes, 7 seconds many ways as we could. You can kind of find this reflected all the way through the stack. Uh just to call out a few of them. Records are encoded as canonical seabore. We inherit that from 17:16 17 minutes, 16 seconds [Music] 17:17 17 minutes, 17 seconds uh the repository is this unisit data structure called an MST. Ununice basically means that uh if you have the same content inside of it, you'll always 17:26 17 minutes, 26 seconds have the same root hash. It's uh has a deterministic construction. even just the fact that we used a repository that could hold all of the data so that you 17:34 17 minutes, 34 seconds know that you have the full view on a user instead of these free floating objects that you have to run around and collect. Uh we have a constrained but 17:41 17 minutes, 41 seconds extendable set of DIDs and hashes. So right now we only support one hash function. We support two DID methods but we use this flexible structure to encode them so that those can expand over time. 17:51 17 minutes, 51 seconds But I think that this reliance on structure is probably best represented by lexicon which is our schema system. 17:57 17 minutes, 57 seconds And we essentially took a schematic approach rather than a semantic approach. Uh and I'll try not to butcher this too much so that the RDF people 18:04 18 minutes, 4 seconds don't get mad at me. Uh but uh a lot of the semantic approaches kind of try to define this universal ontology or interoperation between different apps. 18:14 18 minutes, 14 seconds And here I have like four different possible representations of a post and they all kind of look similar. You can kind of imagine how you can interpret 18:21 18 minutes, 21 seconds them all in the same context. And the question is are they the same thing? Can we interpret them in the same context? 18:27 18 minutes, 27 seconds And instead of trying to build this universal ontology, we tried to give very specific tools for determining if you're talking about the same thing. So 18:35 18 minutes, 35 seconds the approach of app proto was to have schemas. Here we have a simplified version of the schema for a blue sky post. And then three examples on it. And 18:44 18 minutes, 44 seconds you can see the first one passes schema validation. The next one looks nothing like it. So we say that's not the same thing. The third one, this green grass 18:52 18 minutes, 52 seconds post, actually does match the schema, but it declares itself as a different schema. And so we say it doesn't pass schema validation. And what this does is 19:00 19 minutes it gets you out of the gray area. You don't have these questions of are we talking about the same thing? This kind of looks like that other thing. Can we interoperate these? It says black and 19:09 19 minutes, 9 seconds white. Either these things are the same thing and we're on the same page or they're not the same thing. We're not on the same page. If they're not the same thing, an application can still use both 19:17 19 minutes, 17 seconds of them, but at least we know that we're talking about different things in these two different contexts. 19:23 19 minutes, 23 seconds Uh, and then the last sort of posture that I'd call out is what I call lazy trust. And this is this idea that you solve trust through checks and balances 19:32 19 minutes, 32 seconds and credible exit instead of trying to operate completely trustlessly. Uh, this notion of being trustless was something that came out of peer-to-peer movements 19:40 19 minutes, 40 seconds and especially blockchain movements. But you actually get a real benefit from trusting services to operate on your behalf. uh doing trustlessly offloads 19:48 19 minutes, 48 seconds that into compute or time or resources or something like that and things just work better if you trust someone that's doing something for you. But you need to 19:57 19 minutes, 57 seconds have the fallback in case they do you wrong. And also the existence of that fallback uh decreases the incentive for them to do something wrong by you. So 20:05 20 minutes, 5 seconds here we have me and I love my PDS and I love Blue Sky and I trust them and I'm offloading some work to them. Uh the relationship between the PDS and Blue 20:14 20 minutes, 14 seconds Sky operates over the self-certifying protocol. So that part actually is trustless. But since we have a human in the equation here, let's try to save 20:21 20 minutes, 21 seconds some time and resources. Now what happens if my PDS goes evil, right? 20:26 20 minutes, 26 seconds Fortunately, because identity is in primacy and it floats above these PDS's, I can migrate my account and signing key over to that PDS. The benefit of having 20:35 20 minutes, 35 seconds this trust with the PDS is that we can hoist the signing key up from the client into the PDS. And there's a lot of benefits that come along with that. you 20:42 20 minutes, 42 seconds don't have to do all of the horrible key management UX that nobody quite has figured out yet. Uh, and whenever I migrate my account, my signing key is 20:51 20 minutes, 51 seconds living over there and I'm happy and I love and trust my PDS again. Similarly with the uh blue sky app view or the 21:00 21 minutes blue sky app, uh what if it goes evil and starts serving bad views of the data uh or what if it just starts making bad product decisions? 21:09 21 minutes, 9 seconds uh the blue sky app view is operating on the same open data that any other application in the network can run on 21:17 21 minutes, 17 seconds and any other application in the network has access to. So whenever the blue sky app view is serving views of that data, it's sort of staking its reputation on 21:25 21 minutes, 25 seconds every view that it serves and it can be called out on that. Similarly, if it starts making bad product decisions, someone else with access to the same data can start making good product 21:34 21 minutes, 34 seconds decisions and people can start using that application. So, if the blue sky app view does me wrong for whatever reason, uh you know, hopefully people 21:42 21 minutes, 42 seconds here start running a self-hosted uh blue sky app, but people can move over to another application that is using that data in a way that matches with their user expectations. 21:53 21 minutes, 53 seconds Uh so wrapping up uh I kind of already went over the influences from these movements but we uh I would sort of situate at proto at the center of these 22:00 22 minutes three movements influenced by the web uh the peer-to-p peer movement and these large distributed systems that underly a lot of the modern web. The things that 22:08 22 minutes, 8 seconds we introduced to it are identity based authority and this split of generic hosting of canonical data from the rich application semantics on top of it. And 22:17 22 minutes, 17 seconds then we tried to approach a lot of these problems uh airing on the side of structure in the protocol, understanding that that helps facilitate coordination and relying on trust with services but 22:26 22 minutes, 26 seconds always having credible exit from those services and using that as an incentive to keep people playing along with each other. So thank you. 22:39 22 minutes, 39 seconds Uh awesome. Thank you so much, Daniel. All right, let's get to questions. 22:45 22 minutes, 45 seconds Anybody have All right, we'll start with Chris. 22:50 22 minutes, 50 seconds It's kind of a question, but um one of the uh things that we found really common in the early days of the web when we were studying how people were using 22:58 22 minutes, 58 seconds HTML and HTTP to share information and what the browsers were forced to render despite the incorrectness of the schema was also the thing we found that really 23:07 23 minutes, 7 seconds helped the web grow astonishingly fast and in a distributed and in and decentralized manner. Right. So when I 23:14 23 minutes, 14 seconds saw your thing about, oh, this schema is conformant, this one's not, so we can't render it, even though we could if we wanted to. Are you hitting any of these 23:21 23 minutes, 21 seconds parts where you're like, if we just accept the misspelling of the schema, we can get this content up the way that the 23:28 23 minutes, 28 seconds people are intending. It's I'm wondering if you're hitting any of those kinds of funny spots there. Yeah. Yeah. And I mean there always is a tangle there 23:36 23 minutes, 36 seconds because I mean we ran into that for instance with timestamps uh and like some clients were putting up time stamps that didn't match our schema 23:44 23 minutes, 44 seconds and that is really unfortunate to deny those. But the other side of that is if we start accepting those then every single client every single service needs 23:52 23 minutes, 52 seconds to be able to process those timestamps right and so it's this question of okay do we let this this piece of data through that we can kind of understand 24:01 24 minutes, 1 second or do we put the implementation burden on every single other client and every single other service to be able to understand that type of data and what we 24:09 24 minutes, 9 seconds eron is that fail early give people feedback that this does not pass validation and you keep that implementation burden off of all of the downstream services. 24:21 24 minutes, 21 seconds Um, so I'm gonna kind of I want to get your take on the question I asked Nick earlier which was around um 24:29 24 minutes, 29 seconds compositional APIs. Um, so like I totally agree with you like having strict API contracts with data is super 24:37 24 minutes, 37 seconds important and it makes it a lot easier to develop against. Mhm. Um, but then the thing that I've seen in a few 24:47 24 minutes, 47 seconds ecosystems is if you get too strict on like, well, this is what the API is and it has these 24:54 24 minutes, 54 seconds 15 fields, then if you don't satisfy all those 15 fields, you're not talking the same language. And so I'm wondering if 25:03 25 minutes, 3 seconds there was ever a thought around like instead of just having a single schema as a spec having like a list of schema 25:13 25 minutes, 13 seconds objects that represent the data contained in that document. Yeah. Yeah. 25:18 25 minutes, 18 seconds That's a that's a great question and you sort of run into a lot of the same data modeling things that people do uh in like more centralized systems with like 25:26 25 minutes, 26 seconds uh protobuffs for instance where there's kind of this rule of thumb that you shouldn't use required fields because they're hard to evolve. Um, and that's something that we've learned as we're working on our lexicons is that I think 25:34 25 minutes, 34 seconds that we use less required fields than we did early on. We put more stuff in open unions whenever we uh even whenever we don't really need it or just structure 25:43 25 minutes, 43 seconds things in a way that it's more likely to be backwards compatible. Um, I do think open unions are sort of like the they're 25:50 25 minutes, 50 seconds the extension point for lexicons and I think that they give you a lot of freedom for what you're talking about and that's kind of I think what you're getting at is embedding schemas inside 25:58 25 minutes, 58 seconds of uh embedding like maybe smaller schemas inside of the routes that you're talking to. Um, so yeah, I yeah, I definitely agree. 26:08 26 minutes, 8 seconds Hi. Hi. Uh, so I've Oh, yep. Nick Jerichus um stuff. Great. Uh two 26:16 26 minutes, 16 seconds thoughts or two two questions and thoughts I should say. Um first uh how do you what are your tactical and 26:25 26 minutes, 25 seconds strategic plans or thoughts around improving uh the safety of users with 26:32 26 minutes, 32 seconds rogue PDS's? It's like if a PDS decides to, you know, yoink out and delete your stuff, then you're because it has the 26:39 26 minutes, 39 seconds keys to your PLC, uh, did you're, you know, Mhm. you're up a creek. Um, and 26:46 26 minutes, 46 seconds then second, what are your thoughts and what's the calculus around infrastructure that becomes a bad actor, not necessarily the app view or the uh 26:56 26 minutes, 56 seconds PDS or, you know, the the the systems themselves. What if what if for example as riding through North America you know 27:03 27 minutes, 3 seconds two providences have to deal with you know that. Yep. So on the first one with 27:11 27 minutes, 11 seconds rogue PDS's I do think that you take on some responsibility whenever you migrate to a new uh PDS uh whether it's one that 27:19 27 minutes, 19 seconds you run or whether it's one someone else runs. The fortunate thing about at proto is that you can scale up in that responsibility and kind of take it on whenever you're ready for it. Uh, I 27:27 27 minutes, 27 seconds think a lot of this comes down to tooling and having uh good tools for setting up recovery keys, good tools for backing up your repository, for auditing 27:36 27 minutes, 36 seconds the PLC log and getting an email or a push notification or something like that if a uh operation happens on your 27:43 27 minutes, 43 seconds identity. Um, so yeah, I think a lot of it comes down to tooling on that. I think that uh hosts should offer that. I think people interested in the hosting 27:51 27 minutes, 51 seconds space should build stuff like that. I know I want to work on stuff like that. 27:55 27 minutes, 55 seconds Um and but yeah, you do take on some responsibility whenever you migrate your account off and then it's just building tools that make it easier to do that. In 28:03 28 minutes, 3 seconds terms of underlying infrastructure, uh similar to how the web is kind of like predicated on the security model of 28:10 28 minutes, 10 seconds the uh or the uh architecture of the internet like at proto is also predicated on the internet. So to some degree we can't really solve routing 28:19 28 minutes, 19 seconds issues at the internet layer. We do layer on additional self-certifying mechanisms on top of it. So, it's not 28:26 28 minutes, 26 seconds like a uh it's not like an ISP can corrupt data or something, but they can do a denial of service attack and prevent packets from getting through. 28:34 28 minutes, 34 seconds Um, bigger problems to deal with. Yeah, I do think that there are bigger problems, but you know, hopefully someone can set up something that can get the packets through. 28:44 28 minutes, 44 seconds Hey, um do you see it as the responsibility of the protocol to sort of define some semantics for um well you 28:54 28 minutes, 54 seconds mentioned like time stamps for example as one thing that one app view might do one thing and another app view might do another thing or but there's all kinds of things like deletion of records or 29:02 29 minutes, 2 seconds just like what you you it's not entirely predictable just based on what's in the lexicon what um an application might do 29:10 29 minutes, 10 seconds with uh a given like pattern of writes or say, you know, you rewrite a record like does that show up as an edit? Does that show up as like does it get cached? 29:19 29 minutes, 19 seconds Does it get does only the first version get uh stick around? Like do you see those as like conventions that ought to be defined in the protocol or just by 29:28 29 minutes, 28 seconds like app view authors or is this something that you've given a thought to? Uh I think most of those get defined by lexicon authors. There's some 29:35 29 minutes, 35 seconds semantics that are in the protocol. It's like every record has a URI. uh every record can be created, updated or deleted, stuff like that. Um but a lot 29:43 29 minutes, 43 seconds of those application semantics are defined by the lexicon developer and kind of my sense is that you should be able to look at record schemas and the 29:51 29 minutes, 51 seconds methods that come along with it and basically have an idea of how an application is constructed. Um and then also projects like lexicon community and 30:00 30 minutes uh lexhub and stuff can provide further documentation on how you're supposed to process stuff. For instance, in the blue blue sky case, like blocks, what exactly 30:07 30 minutes, 7 seconds are you supposed to do whenever there's a block? Um, providing further documentation on that for applications to build on. 30:16 30 minutes, 16 seconds Hi, uh, I'm Eric. Uh, so I'm just curious, so you're talking a little bit about rogue PDS's, um, credible exit, 30:23 30 minutes, 23 seconds that sort of thing. Uh I'm just curious what your general thoughts are on these are really important things to users in 30:31 30 minutes, 31 seconds particular and they should be things that are better understood by users in particular. Um what are your thoughts on 30:38 30 minutes, 38 seconds enabling users to understand uh the fundamental underlyings of the protocol and the fact that there are 30:45 30 minutes, 45 seconds options available to them when backed actors come into play. uh especially with uh looking into how people might 30:52 30 minutes, 52 seconds consider self-hosting or doing these sorts of things that are potentially more technical for the majority of the user bases of you know whether it be 31:01 31 minutes, 1 second Blue Sky or any emerging platforms on the protocol. Uh so I'm just curious what your general thoughts are on you 31:08 31 minutes, 8 seconds know how can we make people aware that this is how the network works rather than a Twitter 2.0 centralized service. 31:17 31 minutes, 17 seconds Yep. Totally. Uh I mean I think that the the thing that you can't do is force users to understand that, right? You have to give them the ability to 31:26 31 minutes, 26 seconds gradually ramp up. And so it's like you have an account on the protocol, that's great. Okay, you set up a recovery key, even better. You're doing backups of 31:34 31 minutes, 34 seconds your repo, even better. Okay, you're self-hosting your account. Wow, even better. Okay, you're doing that and you're using a self-hosted uh 31:41 31 minutes, 41 seconds application. Awesome. Even better. Uh so you need to give this on-ramp for users to gradually adopt more responsibility 31:49 31 minutes, 49 seconds for their experience in the protocol but you can't thrust all of it on them at the beginning and then it's just a question of uh building good UX around 31:57 31 minutes, 57 seconds that making it approachable figuring out how to explain the ideas. Um I think that there are ways to do it. It is a hard question. I joke that blue sky is a 32:06 32 minutes, 6 seconds like hat that looks like Twitter sitting on top of a you know distributed systems company but like it's it's a skeworphism 32:14 32 minutes, 14 seconds that people can understand and recognize and gradually introduce people to how weird we can get. Yeah. 32:25 32 minutes, 25 seconds Yeah. Go see Miss Boba as well. Hey Daniel. Uh Sam here from Fedica. I have a question about car files. Uh do you 32:33 32 minutes, 33 seconds guys plan to continue to use car files uh in the future? And if so uh because they're a little difficult to work with 32:42 32 minutes, 42 seconds from the developer community and if so do you plan to make it easier just like how Jetream made it simpler than than 32:49 32 minutes, 49 seconds the fire hose? Yeah, I do think that we plan to use car files going forward. Uh we really want to write our own car 32:56 32 minutes, 56 seconds library that has better interfaces. The current one is not very great. Uh so uh yes we do want to use them. Yes we want 33:04 33 minutes, 4 seconds to make them easier to use. And then on your question of Jetream my hope is that we evolve Jetream into a a larger 33:11 33 minutes, 11 seconds broader purpose like sync facilitation service that not only provides this like JSON API of records but also facilitates 33:18 33 minutes, 18 seconds backfill and repo sync and some of these headier sync operations. 33:28 33 minutes, 28 seconds Other questions? Otherwise, I definitely have a question for Daniel. 33:31 33 minutes, 31 seconds Um, the I'm I'm interested in hearing more about how you define the fear and the sort of threat modeling, right? Like 33:39 33 minutes, 39 seconds one of the exciting parts about the approach that y'all are taking as a company is something that um Brian 33:46 33 minutes, 46 seconds Fitzpatrick talked about um with um the data liberation front at Google and the idea that um companies should always 33:54 33 minutes, 54 seconds make it possible for their users to be able to migrate away because that is the way that you can say like your users are affirmatively committed to the decisions 34:03 34 minutes, 3 seconds that you're making and why you're you know that that y'all can maintain alignment and that you're not just trapping them there, right? Um, when it 34:12 34 minutes, 12 seconds comes to how you're thinking about that both at the protocol design level, like how evil are you imagining that your apps can get, like how does that 34:19 34 minutes, 19 seconds actually tangibly manifest in the way that you're thinking and sort of threat modeling the types of bad actions that you could take? Yeah, totally. 34:28 34 minutes, 28 seconds Um, I mean, I don't know how evil they can get. There's messed up people that that's kind of thing, right? Like the the the threat modeling part about this is like, you know, I worked with 34:36 34 minutes, 36 seconds journalists for a really long time and journalists can have any conceivable threat model from like I only work on public records and like I don't care what people see all the way to like if 34:45 34 minutes, 45 seconds anybody finds my files like somebody's going to get their fingernails torn out, right? Um and and so like the the thing that impresses me about the way that you 34:53 34 minutes, 53 seconds all think about design is the decomposition of the problem space and and so uh uh figuring out how to teslate out like which problems you want to 35:01 35 minutes, 1 second threat model I think is an interesting way that you go about this. Yeah, totally. I I think like the big way that I think about it is that 35:09 35 minutes, 9 seconds uh you since you always have that credible exit, it prevents one rentseeking behaviors where people can abuse their position as like some 35:18 35 minutes, 18 seconds central provider and uh it also uh gives an incentive not to abuse users on your service and it also gives the ability to 35:27 35 minutes, 27 seconds people for people to make you know very bespoke services to fit certain needs. 35:31 35 minutes, 31 seconds So if someone has a very high privacy need, they can use a service like that and interoperate with the rest of the network. So to I don't know, maybe to 35:39 35 minutes, 39 seconds compare it to like messaging systems or something, it's like uh you know, my aunt is never going to use signal. I there's nothing I could do to convince 35:46 35 minutes, 46 seconds her to do that. But if I'm very privacy conscious, it would be great if I could use signal and still talk to my aunt. Uh and similarly it's like if someone has 35:55 35 minutes, 55 seconds like very high privacy needs uh with this choice of service uh or high privacy or data security needs with this ability to choose your service hopefully 36:04 36 minutes, 4 seconds you could choose a PDS service provider that is really tailored to those. 36:09 36 minutes, 9 seconds Um so yeah I think that the the the choice of service gives you a lot of optionality on that. Cool. Thanks. All right. Other questions? Otherwise we'll say yeah come on over. 36:24 36 minutes, 24 seconds Proton PDS Daniel from the Bay. So, I was I had a 36:33 36 minutes, 33 seconds conversation about this earlier in the week. Um Boris had chimed in, but um more so a non-technical question. Um how 36:40 36 minutes, 40 seconds has Blue Sky discussed how to signal to people that they can sign in um to like an app for example with their PDS? Like 36:49 36 minutes, 49 seconds let's say for example, you meet someone who doesn't know what Blue Sky is, but there's a they're signing into an app or they're about to they're at the login screen of an app that is on, you know, 36:58 36 minutes, 58 seconds the AT Pro protocol. How do you signal to them that they can s that they can sign into um sign into that app because 37:06 37 minutes, 6 seconds you can't say sign into Blue Sky? Yeah, totally. What is that, right? You know, that's a great question. That's a UX thing we need to figure out. If people 37:13 37 minutes, 13 seconds have good ideas on this, I definitely want to talk with them about it. But uh the great thing about app proto identity is that it can serve as this universal 37:21 37 minutes, 21 seconds identifier both for social apps that run on top of the app protocol but also for non uh app protocol apps and uh figuring 37:29 37 minutes, 29 seconds out some UX pattern to be able to posit it like sign in with atmosphere but that's kind of wordy and confusing. So we got to figure out a good one. 37:37 37 minutes, 37 seconds Yes, sign in with cloudy ad symbol because in the GitHub discussion, I know there was a specific mention that um I 37:46 37 minutes, 46 seconds believe it was Jay that um mentioned that she didn't want at Proto actually used as like for branding. 37:56 37 minutes, 56 seconds Some someone in the someone mentioned it's either someone mentioned that um leader said that um she didn't want at 38:04 38 minutes, 4 seconds Proto used for branding or that branding um that at Proto shouldn't be used for branding. So I was just like well what 38:12 38 minutes, 12 seconds what should we do you know? Yeah I'm actually not quite sure on the context of that. So maybe maybe we can chat about it after. But uh we definitely do 38:20 38 minutes, 20 seconds want people to understand the atmosphere and have a coherent conception of the atmosphere and know that you can log into stuff with the atmosphere. And I do think a lot of that is going to be the brand and apps being able to use that 38:29 38 minutes, 29 seconds brand. This is also one of the things that came up in one of the tech talks that uh Boris hosted a couple of months ago. And I I definitely also see that this is a community responsibility in 38:38 38 minutes, 38 seconds part and this is one of those things that we can work we need to be able to work on together. But also you'll have the largest install base, right? So, yes, we're interested in engaging y'all 38:46 38 minutes, 46 seconds on that more. Um, all right. Uh, I think we've got time for one more question if anybody's got All right. 38:55 38 minutes, 55 seconds Okay, we can do two. 39:00 39 minutes Um, hi. Yeah. Uh so I think it's in the lexicon community and maybe Nick can correct me but there is a github discussion on this topic you know how to 39:09 39 minutes, 9 seconds uh improve the UX and the UI flow of logging through other methods. Uh I forgot the link but I think it should be 39:17 39 minutes, 17 seconds under one discussions on the lexicon community GitHub repo. Cool. Thanks. 39:25 39 minutes, 25 seconds Hi. There was a long list of like previous references that informed the atproto ethos. I'm curious if there are protocols out there today that you think 39:34 39 minutes, 34 seconds are at ProtoE or share similar values or just things like that. 39:39 39 minutes, 39 seconds [Music] 39:42 39 minutes, 42 seconds Um, I'd say surprisingly Noir actually has a similar architecture to app proto in terms of having uh relays 39:50 39 minutes, 50 seconds broadcasting stuff out and aggregating it in applications although they don't really have the uh like very rich 39:58 39 minutes, 58 seconds backends that uh app proto has with app views. Um, but it sort of has a similar shape of information flow. 40:08 40 minutes, 8 seconds Um I don't know the web still. 40:14 40 minutes, 14 seconds Yeah. Um thank you very much. Uh this is great. Uh appreciate all the questions as well. Sweet. Thanks guys.