Something went wrong. Try again.
about things notes.zzstoatzz.io
notes
Something went wrong. Try again.
Markdown
the way that it does.
0:02
2 seconds
Sweet. Can everybody hear me? Okay, good stuff. Um, cool. I'm going to be talking about appro ethos, which is
0:09
9 seconds
essentially the philosophical and aesthetic principles underlying the design of the protocol. So, before we get started, my name is Daniel Homegrren. I'm the head
0:18
18 seconds
of protocol here at Blue Sky. Uh, you can find me on the atmosphere at dhomes.xyz. And for the freaks out
0:26
26 seconds
there, you can find me at this did So the title for this talk uh actually
0:33
33 seconds
came from a thought that I have had whenever I saw Martin post this. He said whether something is decentralized or not is a function of the administrative
0:41
41 seconds
control of different parts of the system not a function of the network topology.
0:45
45 seconds
And this just struck me so hard and I was like yes this is exactly what we're going for. I actually brought this up to Martin a couple of days ago. He didn't even remember posting it. I'm like I think about this every week Martin.
0:58
58 seconds
uh but I quoted this and I said this is appro ethos and uh I do think that this is very descriptive of the topology and
1:06
1 minute, 6 seconds
the power dynamics in the network uh but what it doesn't quite have is like a prescription for how to get there and so this term ad protoethos kept coming up
1:14
1 minute, 14 seconds
in our protocol design questions we kept kind of like nice uh we kind of kept coming back to it whenever we were weighing different
1:22
1 minute, 22 seconds
options and I wanted to take the chance to try to dissect what that ethos is the historical influences on the design of the protocol, some of the new things
1:31
1 minute, 31 seconds
that we brought to it, as well as some as well as some of the postures that we adopt as we are working on the design of it. Uh, and just for what it's worth,
1:39
1 minute, 39 seconds
not all of this was like perfectly packaged or anything from the beginning.
1:42
1 minute, 42 seconds
We had like a million false paths and dead ends and hard debates between team members and everything. And maybe in a future talk, we can give the philosophical history of app proto, but
1:51
1 minute, 51 seconds
I'm going to give you guys the nice clean packaged version of it.
1:55
1 minute, 55 seconds
So I would sort of situate ad proto at the center of uh three movements or three trends and influenced by those three trends. On the one hand we have
2:03
2 minutes, 3 seconds
the web. Uh then we have the peer-to-peer movement and we have I don't have a good quick word for this but kind of large distributed systems or
2:11
2 minutes, 11 seconds
these data inensive applications that kind of run the modern web. I also didn't have a good picture for that so I used the bore from Martin's book. Uh so
2:19
2 minutes, 19 seconds
we'll start with the most obvious influence which is the web. Uh, as I'm sure you guys all know, the web is an information system for publishing documents over the internet. So, if you
2:26
2 minutes, 26 seconds
have a static IP address and you have a server, you can host a document. If you have an internet connection and a route to that server, you can fetch that document. Uh, kind of the incredible
2:35
2 minutes, 35 seconds
thing about the web, the defining feature of it is that it's completely permissionless. That's the only thing that you need to publish a document.
2:40
2 minutes, 40 seconds
That's the only thing that you need to go fetch a document. Location is also based in the authority. We we layer on DNS on top of those IP addresses. This
2:48
2 minutes, 48 seconds
sounds really obvious to say out loud, but you know that a document is the thing that you get from a server because you talked to that server and the server gave you that document, right? But the
2:57
2 minutes, 57 seconds
authority rests in the location that you talk to whenever you're fetching the document. Uh this combined with the permission list of the web was actually
3:04
3 minutes, 4 seconds
a pretty radical idea. You didn't need any central publisher, any central distributor. You didn't need a certification checklist to say that you
3:11
3 minutes, 11 seconds
could publish a document or anything like that. If that permissionless and openness was the defining governance feature of the web, sort of the defining
3:20
3 minutes, 20 seconds
product feature of the web was the hyperlink which gave you interconnected data. Uh documents could link to other documents and specifically they could
3:28
3 minutes, 28 seconds
link across authority again completely permissionlessly. The downside of all of this is that servers are hard. These all run on servers. Not everybody knows how
3:36
3 minutes, 36 seconds
to run a server. Not everyone can get a static IP. Not everybody can configure DNS. And so what we started to see emerge in the modern web or what people
3:43
3 minutes, 43 seconds
call web 2.0 is this move from permissionless publishing to platforms.
3:48
3 minutes, 48 seconds
Everything running on a big server called bigocial.com which I thought about making a BS joke but I thought that might be dangerous with the name of our
3:57
3 minutes, 57 seconds
company. Uh so these these platforms were account based. You'd register an account with one of these platforms. The benefits of it is that it's very easy to
4:05
4 minutes, 5 seconds
post. You don't need any technical knowledge. You just need an account on the platform. you have rich media, uh, photos, videos, music, interactions,
4:13
4 minutes, 13 seconds
dynamic real-time interactions between accounts on that data. And then to top it all off, a nice algorithm that churns through all of that data, uh, shows you
4:21
4 minutes, 21 seconds
exactly what you want to see, crunches all those like counts, the reply threads, everything like that. Uh, and I don't want anyone to think that I'm like
4:29
4 minutes, 29 seconds
doy nostalgic for web 1.0. There's a lot about the modern web that is really great. Nobody wants to go away from that, but there was also something that
4:36
4 minutes, 36 seconds
was lost with it. So that kind of takes me to the next influence on at proto which is the peer-to-peer movement. In a lot of ways this was a response to what
4:44
4 minutes, 44 seconds
was happening in the modern web and that sort of centralization in these large social platforms. And the basic question that it asked is why is there a server
4:53
4 minutes, 53 seconds
in the mix? So I have me here and Devon over there. I'm writing a post. Devon follows me. I want to send that post to Devon. And we say why is the server
5:00
5 minutes
here? What if we just removed it and I sent the post directly to Devon from my phone to his phone? And the question that came out of that is can you build a modern social network without a backend?
5:10
5 minutes, 10 seconds
Uh this kind of came from a distrust of these large centralized uh server services and platforms. And so if I if
5:17
5 minutes, 17 seconds
uh Paul, Devon, and Brian are all following me, I make a post. I send that post out to the three of them directly from my device to their device. Of
5:25
5 minutes, 25 seconds
course, there's a problem there. If I'm not online at the same time as them, how does my post actually reach out to them?
5:31
5 minutes, 31 seconds
And so these peer-to-p peer pro uh protocols said,"Well, if I sent my post to Paul already and I'm offline and he's online at the same time as Devon and
5:38
5 minutes, 38 seconds
Brian, then he can forward on the post to them." Of course, the problem with that is that what if Paul replaces my post with a poop emoji and sends that to
5:46
5 minutes, 46 seconds
Devon, and that's really embarrassing for me. I don't want Devon to think that I posted that. And so these protocols would get around that by introducing public key cryptography signing the
5:56
5 minutes, 56 seconds
message. Uh which essentially put the certification that the message is what I intended it to be in the message. Every
6:03
6 minutes, 3 seconds
account would have a public key associated with it and it gave me the liberty to broadcast my message out and trust that other people could deliver it reliably with all without altering the
6:11
6 minutes, 11 seconds
contents of it. Uh then in a lot of these protocols you would then get content address data.
6:18
6 minutes, 18 seconds
uh you would sort of build this abstraction around the routing of these messages. Since you could get them from anywhere, the certification is in the data itself. You would reach out and you
6:27
6 minutes, 27 seconds
would say, I know that there's this piece of data with this hash and you would reach out to the network to go fetch that and grab it from some device in the network. What this gave you
6:35
6 minutes, 35 seconds
whenever it was functioning well is this really beautiful abstraction. It's what drew me to the peer-to-peer movement in the first place, which is it completely dissolves away the network topology, the
6:44
6 minutes, 44 seconds
devices, the locations, everything like that. And there's like kind of this essential essence to these posts that are kind and you can just reach into the
6:52
6 minutes, 52 seconds
ether and pull them out. It's like a really beautiful thing. It's a really beautiful concept whenever it's working well. Of course, that's if the abstraction is working. A lot of the times it doesn't work well and you have
7:01
7 minutes, 1 second
uh data availability problems, data routing problems, uh things like that. And then the big problem is that phones are a lot
7:09
7 minutes, 9 seconds
smaller than data centers. And we have these big machines in the cloud that are crunching all of this data for us. And phones just cannot keep up with that same data. you have resource usage
7:17
7 minutes, 17 seconds
problems, battery life problems, and then you just can't store all of that data on your phone. So, you're not going to be able to create the same global views that these big social platforms
7:25
7 minutes, 25 seconds
are able to create. Um, excuse me.
7:32
7 minutes, 32 seconds
So that brings us to the last influence on approto which is the parallel development that was going on while these peer-to-p peer protocols were gaining or were kind of being um worked
7:41
7 minutes, 41 seconds
on which is these how these large distributed systems that underlied the modern web were being architected. So
7:50
7 minutes, 50 seconds
sort of the features of these is that they have millions and eventually billions of users uh real-time dynamic views. Someone likes a post you get an update on that milliseconds after the
7:58
7 minutes, 58 seconds
fact. They're very very readheavy. For every post that is posted, it might get loaded thousands or millions of times.
8:04
8 minutes, 4 seconds
And there's very high expectations whenever it comes to downtime and latency, especially as more and more of the interaction on the web started centralizing on these
8:12
8 minutes, 12 seconds
platforms. So, we needed new techniques for building these data inensive applications. Uh Martin Kleman's book designing data intensive applications is
8:20
8 minutes, 20 seconds
kind of the bible on this. He gave a really great overview of all of the trends. Uh I'm going to talk about kind of high level the trends. Obviously, all
8:28
8 minutes, 28 seconds
of these platforms are architected uh differently, but maybe we can sus out some of the commonalities in a lot of them. So, maybe one of these platforms
8:35
8 minutes, 35 seconds
starts out with a server and a big Postgress database. Uh, and the first thing that you're going to run into because they're so right read heavy is
8:43
8 minutes, 43 seconds
that you want to split out the read and write loads. Uh, so here we see that database getting some read replicas. All of the writes are still coming in on a
8:51
8 minutes, 51 seconds
leader database. The reads are getting pushed out to these read replicas. you can scale up reading separately from the rights that are coming
8:58
8 minutes, 58 seconds
in. A lot of these platforms started to decompose their monolithic services into microservices. You would essentially peel off some piece of the service uh
9:07
9 minutes, 7 seconds
wrap it in its own service. The monolith or the main service would then call into it. There are a lot of reasons to do this. Organizational reasons. You could
9:15
9 minutes, 15 seconds
have small teams own code bases. Uh you could also scale resource usage according to what that service did.
9:22
9 minutes, 22 seconds
video transcoding looks very different from timeline generation and you could put these on machines and give them resources that matched what they needed
9:28
9 minutes, 28 seconds
to do. Um, of course, whenever we do this, all of a sudden we just created a distributed system. And whenever you
9:36
9 minutes, 36 seconds
have a distributed system, you have distributed problems. Uh, you no longer have strong consistency necessarily between all of these services. If one of
9:44
9 minutes, 44 seconds
these microservices goes down for three hours and it comes back up, did it just miss out on three hours of writes? Is it ever going to get back in contact with
9:51
9 minutes, 51 seconds
the same data that the other services in the data center have? And so we find these services starting to lean into eventual consistency uh and oftentimes
10:00
10 minutes
eventually or oftent times uh using eventually consistent databases as well.
10:05
10 minutes, 5 seconds
I put Sila here because that's what we use at Blue Sky. Uh but it's a successor to Cassandra. There's other data databases of that type as well. And
10:13
10 minutes, 13 seconds
these kind of give up on the acid guarantees of Postgress or something like that. give up on the ability to uh linearize writes and uh really focus on
10:22
10 minutes, 22 seconds
being able to do low latency writes and reads. Another aspect of these is that you also give up on the ability to dynamically query your data. You kind of
10:30
10 minutes, 30 seconds
have to be aware of what your application is going to ask of the data uh before you lay it out in your database. That means that if you need a new index or you need to reindex stuff, you have to reprocess all of the data.
10:41
10 minutes, 41 seconds
If the data is already in the database, how are you going to do that? A lot of these problems were generally solved by an approach of stream processing. I
10:49
10 minutes, 49 seconds
called out Kafka here, but again there's other examples of uh that. And this kind of introduced a information flow into these backends where writes would come
10:56
10 minutes, 56 seconds
in, they would get added to uh the stream processor and then filter out to the microservices or the database. This meant that if a microser went down for
11:05
11 minutes, 5 seconds
three hours, it could pick back up, talk to this canonical stream of data and catch back up, know that it has everything. Similarly, if you have to
11:12
11 minutes, 12 seconds
rebuild an index or build a new index, you can just replay the entire stream of canonical data and know that you have everything in that index. These kind of
11:20
11 minutes, 20 seconds
these canonical stream processors sort of served as the backbone to these data centers uh that could uh give the
11:27
11 minutes, 27 seconds
feeling of consistency in them. So, this brings me back to approto ethos. Uh I would say that at proto ethos is sort of
11:35
11 minutes, 35 seconds
situated as the synthesis of these three movements. Uh so it's influenced by all of these movements. From the web, we kind of took the fact that it's open,
11:43
11 minutes, 43 seconds
permissionless, and it's this universal uh network of interconnected data. From the peer-to-peer movement, uh we took
11:51
11 minutes, 51 seconds
location independence of data, the fact that the data is self-certifying. It has the certification embedded in it and this general skepticism of
11:59
11 minutes, 59 seconds
centralization. And then from these large data inensive distributed systems, the splitting of read and write load uh
12:06
12 minutes, 6 seconds
large application aware indices decomposition of stuff into microservices owning up to eventual consistency and uh stream
12:14
12 minutes, 14 seconds
processing. So now I'm going to get into the uh new things that appro brought to the table. The first is identity based authority. Again, this is in contrast to
12:23
12 minutes, 23 seconds
the location-based authority of the web or the contentbased authority of peerto-peer. We said identity is in privacy. So the central thing is my
12:31
12 minutes, 31 seconds
account at dhomes.xyz and then the thing that you'll find there is all of the social data contained inside of it. With this
12:39
12 minutes, 39 seconds
came a change in addressing scheme. So in the web we have https. It points to a location bigo.com and then you fetch my
12:46
12 minutes, 46 seconds
profile from that. A peer-to-peer protocol like ipfs points to a uh hash of some data. In app proto it points to
12:54
12 minutes, 54 seconds
dhomes.xyz XYZ which kind of looks like a location but it's actually an identity and then it points to the uh record type the collection and the key of the record
13:03
13 minutes, 3 seconds
inside of that. Similar to peer-to-peer protocols this gives you location independent data. Uh every identity even if it looks
13:10
13 minutes, 10 seconds
like a DNS name resolves to a universal identifier that is independent of location and then that identifier in
13:18
13 minutes, 18 seconds
turn resolves to the location where the data is stored and key material. So we essentially uh introduce this layer of indirection between the identity and the
13:26
13 minutes, 26 seconds
location of the data and that lets these identities kind of float on top of the data abstracted away from the actual architect or the actual topology of the
13:35
13 minutes, 35 seconds
network. Uh because they're floating on top of the network they're able to migrate between locations that actually host the data or host that signing key.
13:44
13 minutes, 44 seconds
Um and uh identity is really the primary thing that floats at top this layer of hosts.
13:51
13 minutes, 51 seconds
And then from the web we kept interconnected data but and similar to how the web has interconnected data across authorities at proto also has
13:59
13 minutes, 59 seconds
interconnected data across authorities but those authorities are identities and the data is structured uh schematic data instead of document
14:08
14 minutes, 8 seconds
data. So the next thing that we introduced is this notion of generic hosting and separating application semantics out from canonical data
14:17
14 minutes, 17 seconds
hosting. So this is my identity that we saw earlier. It has some likes, follows, posts, and images inside of it. From the
14:24
14 minutes, 24 seconds
perspective of my data hosting server, my PDS, it just looks like this. They're just Seabore objects. It has no idea what the rich application semantics
14:32
14 minutes, 32 seconds
inside of my repository are. It just goes, this is this is Seabore. What this lets happen is that application developers can then own
14:40
14 minutes, 40 seconds
their schematic space without coordinating with the host of that canonical data. So inside my repository
14:47
14 minutes, 47 seconds
we have uh blue sky posts and smoke signal events and whitew blog entries that are defined by the application
14:56
14 minutes, 56 seconds
app.bvents.smokes signal com.whitewwind not by the host that is hosting that canonical data. All of these uh hosts of data then
15:05
15 minutes, 5 seconds
exist in this very crawable, very accessible network of open data and applications can then reach into that data, pull it out and imbue it with the
15:14
15 minutes, 14 seconds
rich application semantics that you associate with the modern web. So what this gives us is decentralized hosting and centralized
15:22
15 minutes, 22 seconds
product development. The hosts that make up this decentralized network are blissfully unaware of all of the rich applications that are running on top of
15:30
15 minutes, 30 seconds
it. But the rich applications can make choices that let them move like how you expect products uh to move. They can have uh you kind of need to be able to
15:39
15 minutes, 39 seconds
have leadership whenever you're making a product. You need to be able to rapidly iterate. You need to be able to use the tools that you want to use, but you inherit all of the benefits of the
15:47
15 minutes, 47 seconds
decentralized hosting and the decentralized data where it can be extensible, remixable, reusable, accessible to other applications while
15:55
15 minutes, 55 seconds
retaining the benefits of having a uh a uh clear leader for a given product. So the next two things that I
16:04
16 minutes, 4 seconds
want to go over are less like features or something that we introduced and more kind of postures that we took as we were approaching these problems. Uh the first
16:11
16 minutes, 11 seconds
is kind of the acknowledgement that having structure actually gives you freedom in one of these open and peri per permissionless networks. So approto
16:19
16 minutes, 19 seconds
is a multi-party lowcoordination network. You have all of these devices and services talking to one another and the question is how do they all talk to
16:27
16 minutes, 27 seconds
each other and how do you come up to a coherent hole. You want to avoid the tyranny of structurelessness which is essentially this collapse whenever you say hey it's really great we have
16:36
16 minutes, 36 seconds
freedom. we can do whatever we want and you have this collapse in coordination where nobody can actually get anything done because there's no clear leader to actually get that done. Uh we want to
16:44
16 minutes, 44 seconds
spend our time building applications and doing cool stuff and not just focusing on facilitating interoperation, fixing weird edge cases, building defenses
16:52
16 minutes, 52 seconds
against bad actors or trying to coordinate evolution whenever none of us have a uh clear leader that can do that coordination. So what you'll find is
17:00
17 minutes
that at proto is a very structured protocol. We tried to put uh we tried to reduce the number of choices that you could make in the important areas in as
17:07
17 minutes, 7 seconds
many ways as we could. You can kind of find this reflected all the way through the stack. Uh just to call out a few of them. Records are encoded as canonical seabore. We inherit that from
17:16
17 minutes, 16 seconds
[Music]
17:17
17 minutes, 17 seconds
uh the repository is this unisit data structure called an MST. Ununice basically means that uh if you have the same content inside of it, you'll always
17:26
17 minutes, 26 seconds
have the same root hash. It's uh has a deterministic construction. even just the fact that we used a repository that could hold all of the data so that you
17:34
17 minutes, 34 seconds
know that you have the full view on a user instead of these free floating objects that you have to run around and collect. Uh we have a constrained but
17:41
17 minutes, 41 seconds
extendable set of DIDs and hashes. So right now we only support one hash function. We support two DID methods but we use this flexible structure to encode them so that those can expand over time.
17:51
17 minutes, 51 seconds
But I think that this reliance on structure is probably best represented by lexicon which is our schema system.
17:57
17 minutes, 57 seconds
And we essentially took a schematic approach rather than a semantic approach. Uh and I'll try not to butcher this too much so that the RDF people
18:04
18 minutes, 4 seconds
don't get mad at me. Uh but uh a lot of the semantic approaches kind of try to define this universal ontology or interoperation between different apps.
18:14
18 minutes, 14 seconds
And here I have like four different possible representations of a post and they all kind of look similar. You can kind of imagine how you can interpret
18:21
18 minutes, 21 seconds
them all in the same context. And the question is are they the same thing? Can we interpret them in the same context?
18:27
18 minutes, 27 seconds
And instead of trying to build this universal ontology, we tried to give very specific tools for determining if you're talking about the same thing. So
18:35
18 minutes, 35 seconds
the approach of app proto was to have schemas. Here we have a simplified version of the schema for a blue sky post. And then three examples on it. And
18:44
18 minutes, 44 seconds
you can see the first one passes schema validation. The next one looks nothing like it. So we say that's not the same thing. The third one, this green grass
18:52
18 minutes, 52 seconds
post, actually does match the schema, but it declares itself as a different schema. And so we say it doesn't pass schema validation. And what this does is
19:00
19 minutes
it gets you out of the gray area. You don't have these questions of are we talking about the same thing? This kind of looks like that other thing. Can we interoperate these? It says black and
19:09
19 minutes, 9 seconds
white. Either these things are the same thing and we're on the same page or they're not the same thing. We're not on the same page. If they're not the same thing, an application can still use both
19:17
19 minutes, 17 seconds
of them, but at least we know that we're talking about different things in these two different contexts.
19:23
19 minutes, 23 seconds
Uh, and then the last sort of posture that I'd call out is what I call lazy trust. And this is this idea that you solve trust through checks and balances
19:32
19 minutes, 32 seconds
and credible exit instead of trying to operate completely trustlessly. Uh, this notion of being trustless was something that came out of peer-to-peer movements
19:40
19 minutes, 40 seconds
and especially blockchain movements. But you actually get a real benefit from trusting services to operate on your behalf. uh doing trustlessly offloads
19:48
19 minutes, 48 seconds
that into compute or time or resources or something like that and things just work better if you trust someone that's doing something for you. But you need to
19:57
19 minutes, 57 seconds
have the fallback in case they do you wrong. And also the existence of that fallback uh decreases the incentive for them to do something wrong by you. So
20:05
20 minutes, 5 seconds
here we have me and I love my PDS and I love Blue Sky and I trust them and I'm offloading some work to them. Uh the relationship between the PDS and Blue
20:14
20 minutes, 14 seconds
Sky operates over the self-certifying protocol. So that part actually is trustless. But since we have a human in the equation here, let's try to save
20:21
20 minutes, 21 seconds
some time and resources. Now what happens if my PDS goes evil, right?
20:26
20 minutes, 26 seconds
Fortunately, because identity is in primacy and it floats above these PDS's, I can migrate my account and signing key over to that PDS. The benefit of having
20:35
20 minutes, 35 seconds
this trust with the PDS is that we can hoist the signing key up from the client into the PDS. And there's a lot of benefits that come along with that. you
20:42
20 minutes, 42 seconds
don't have to do all of the horrible key management UX that nobody quite has figured out yet. Uh, and whenever I migrate my account, my signing key is
20:51
20 minutes, 51 seconds
living over there and I'm happy and I love and trust my PDS again. Similarly with the uh blue sky app view or the
21:00
21 minutes
blue sky app, uh what if it goes evil and starts serving bad views of the data uh or what if it just starts making bad product decisions?
21:09
21 minutes, 9 seconds
uh the blue sky app view is operating on the same open data that any other application in the network can run on
21:17
21 minutes, 17 seconds
and any other application in the network has access to. So whenever the blue sky app view is serving views of that data, it's sort of staking its reputation on
21:25
21 minutes, 25 seconds
every view that it serves and it can be called out on that. Similarly, if it starts making bad product decisions, someone else with access to the same data can start making good product
21:34
21 minutes, 34 seconds
decisions and people can start using that application. So, if the blue sky app view does me wrong for whatever reason, uh you know, hopefully people
21:42
21 minutes, 42 seconds
here start running a self-hosted uh blue sky app, but people can move over to another application that is using that data in a way that matches with their user expectations.
21:53
21 minutes, 53 seconds
Uh so wrapping up uh I kind of already went over the influences from these movements but we uh I would sort of situate at proto at the center of these
22:00
22 minutes
three movements influenced by the web uh the peer-to-p peer movement and these large distributed systems that underly a lot of the modern web. The things that
22:08
22 minutes, 8 seconds
we introduced to it are identity based authority and this split of generic hosting of canonical data from the rich application semantics on top of it. And
22:17
22 minutes, 17 seconds
then we tried to approach a lot of these problems uh airing on the side of structure in the protocol, understanding that that helps facilitate coordination and relying on trust with services but
22:26
22 minutes, 26 seconds
always having credible exit from those services and using that as an incentive to keep people playing along with each other. So thank you.
22:39
22 minutes, 39 seconds
Uh awesome. Thank you so much, Daniel. All right, let's get to questions.
22:45
22 minutes, 45 seconds
Anybody have All right, we'll start with Chris.
22:50
22 minutes, 50 seconds
It's kind of a question, but um one of the uh things that we found really common in the early days of the web when we were studying how people were using
22:58
22 minutes, 58 seconds
HTML and HTTP to share information and what the browsers were forced to render despite the incorrectness of the schema was also the thing we found that really
23:07
23 minutes, 7 seconds
helped the web grow astonishingly fast and in a distributed and in and decentralized manner. Right. So when I
23:14
23 minutes, 14 seconds
saw your thing about, oh, this schema is conformant, this one's not, so we can't render it, even though we could if we wanted to. Are you hitting any of these
23:21
23 minutes, 21 seconds
parts where you're like, if we just accept the misspelling of the schema, we can get this content up the way that the
23:28
23 minutes, 28 seconds
people are intending. It's I'm wondering if you're hitting any of those kinds of funny spots there. Yeah. Yeah. And I mean there always is a tangle there
23:36
23 minutes, 36 seconds
because I mean we ran into that for instance with timestamps uh and like some clients were putting up time stamps that didn't match our schema
23:44
23 minutes, 44 seconds
and that is really unfortunate to deny those. But the other side of that is if we start accepting those then every single client every single service needs
23:52
23 minutes, 52 seconds
to be able to process those timestamps right and so it's this question of okay do we let this this piece of data through that we can kind of understand
24:01
24 minutes, 1 second
or do we put the implementation burden on every single other client and every single other service to be able to understand that type of data and what we
24:09
24 minutes, 9 seconds
eron is that fail early give people feedback that this does not pass validation and you keep that implementation burden off of all of the downstream services.
24:21
24 minutes, 21 seconds
Um, so I'm gonna kind of I want to get your take on the question I asked Nick earlier which was around um
24:29
24 minutes, 29 seconds
compositional APIs. Um, so like I totally agree with you like having strict API contracts with data is super
24:37
24 minutes, 37 seconds
important and it makes it a lot easier to develop against. Mhm. Um, but then the thing that I've seen in a few
24:47
24 minutes, 47 seconds
ecosystems is if you get too strict on like, well, this is what the API is and it has these
24:54
24 minutes, 54 seconds
15 fields, then if you don't satisfy all those 15 fields, you're not talking the same language. And so I'm wondering if
25:03
25 minutes, 3 seconds
there was ever a thought around like instead of just having a single schema as a spec having like a list of schema
25:13
25 minutes, 13 seconds
objects that represent the data contained in that document. Yeah. Yeah.
25:18
25 minutes, 18 seconds
That's a that's a great question and you sort of run into a lot of the same data modeling things that people do uh in like more centralized systems with like
25:26
25 minutes, 26 seconds
uh protobuffs for instance where there's kind of this rule of thumb that you shouldn't use required fields because they're hard to evolve. Um, and that's something that we've learned as we're working on our lexicons is that I think
25:34
25 minutes, 34 seconds
that we use less required fields than we did early on. We put more stuff in open unions whenever we uh even whenever we don't really need it or just structure
25:43
25 minutes, 43 seconds
things in a way that it's more likely to be backwards compatible. Um, I do think open unions are sort of like the they're
25:50
25 minutes, 50 seconds
the extension point for lexicons and I think that they give you a lot of freedom for what you're talking about and that's kind of I think what you're getting at is embedding schemas inside
25:58
25 minutes, 58 seconds
of uh embedding like maybe smaller schemas inside of the routes that you're talking to. Um, so yeah, I yeah, I definitely agree.
26:08
26 minutes, 8 seconds
Hi. Hi. Uh, so I've Oh, yep. Nick Jerichus um stuff. Great. Uh two
26:16
26 minutes, 16 seconds
thoughts or two two questions and thoughts I should say. Um first uh how do you what are your tactical and
26:25
26 minutes, 25 seconds
strategic plans or thoughts around improving uh the safety of users with
26:32
26 minutes, 32 seconds
rogue PDS's? It's like if a PDS decides to, you know, yoink out and delete your stuff, then you're because it has the
26:39
26 minutes, 39 seconds
keys to your PLC, uh, did you're, you know, Mhm. you're up a creek. Um, and
26:46
26 minutes, 46 seconds
then second, what are your thoughts and what's the calculus around infrastructure that becomes a bad actor, not necessarily the app view or the uh
26:56
26 minutes, 56 seconds
PDS or, you know, the the the systems themselves. What if what if for example as riding through North America you know
27:03
27 minutes, 3 seconds
two providences have to deal with you know that. Yep. So on the first one with
27:11
27 minutes, 11 seconds
rogue PDS's I do think that you take on some responsibility whenever you migrate to a new uh PDS uh whether it's one that
27:19
27 minutes, 19 seconds
you run or whether it's one someone else runs. The fortunate thing about at proto is that you can scale up in that responsibility and kind of take it on whenever you're ready for it. Uh, I
27:27
27 minutes, 27 seconds
think a lot of this comes down to tooling and having uh good tools for setting up recovery keys, good tools for backing up your repository, for auditing
27:36
27 minutes, 36 seconds
the PLC log and getting an email or a push notification or something like that if a uh operation happens on your
27:43
27 minutes, 43 seconds
identity. Um, so yeah, I think a lot of it comes down to tooling on that. I think that uh hosts should offer that. I think people interested in the hosting
27:51
27 minutes, 51 seconds
space should build stuff like that. I know I want to work on stuff like that.
27:55
27 minutes, 55 seconds
Um and but yeah, you do take on some responsibility whenever you migrate your account off and then it's just building tools that make it easier to do that. In
28:03
28 minutes, 3 seconds
terms of underlying infrastructure, uh similar to how the web is kind of like predicated on the security model of
28:10
28 minutes, 10 seconds
the uh or the uh architecture of the internet like at proto is also predicated on the internet. So to some degree we can't really solve routing
28:19
28 minutes, 19 seconds
issues at the internet layer. We do layer on additional self-certifying mechanisms on top of it. So, it's not
28:26
28 minutes, 26 seconds
like a uh it's not like an ISP can corrupt data or something, but they can do a denial of service attack and prevent packets from getting through.
28:34
28 minutes, 34 seconds
Um, bigger problems to deal with. Yeah, I do think that there are bigger problems, but you know, hopefully someone can set up something that can get the packets through.
28:44
28 minutes, 44 seconds
Hey, um do you see it as the responsibility of the protocol to sort of define some semantics for um well you
28:54
28 minutes, 54 seconds
mentioned like time stamps for example as one thing that one app view might do one thing and another app view might do another thing or but there's all kinds of things like deletion of records or
29:02
29 minutes, 2 seconds
just like what you you it's not entirely predictable just based on what's in the lexicon what um an application might do
29:10
29 minutes, 10 seconds
with uh a given like pattern of writes or say, you know, you rewrite a record like does that show up as an edit? Does that show up as like does it get cached?
29:19
29 minutes, 19 seconds
Does it get does only the first version get uh stick around? Like do you see those as like conventions that ought to be defined in the protocol or just by
29:28
29 minutes, 28 seconds
like app view authors or is this something that you've given a thought to? Uh I think most of those get defined by lexicon authors. There's some
29:35
29 minutes, 35 seconds
semantics that are in the protocol. It's like every record has a URI. uh every record can be created, updated or deleted, stuff like that. Um but a lot
29:43
29 minutes, 43 seconds
of those application semantics are defined by the lexicon developer and kind of my sense is that you should be able to look at record schemas and the
29:51
29 minutes, 51 seconds
methods that come along with it and basically have an idea of how an application is constructed. Um and then also projects like lexicon community and
30:00
30 minutes
uh lexhub and stuff can provide further documentation on how you're supposed to process stuff. For instance, in the blue blue sky case, like blocks, what exactly
30:07
30 minutes, 7 seconds
are you supposed to do whenever there's a block? Um, providing further documentation on that for applications to build on.
30:16
30 minutes, 16 seconds
Hi, uh, I'm Eric. Uh, so I'm just curious, so you're talking a little bit about rogue PDS's, um, credible exit,
30:23
30 minutes, 23 seconds
that sort of thing. Uh I'm just curious what your general thoughts are on these are really important things to users in
30:31
30 minutes, 31 seconds
particular and they should be things that are better understood by users in particular. Um what are your thoughts on
30:38
30 minutes, 38 seconds
enabling users to understand uh the fundamental underlyings of the protocol and the fact that there are
30:45
30 minutes, 45 seconds
options available to them when backed actors come into play. uh especially with uh looking into how people might
30:52
30 minutes, 52 seconds
consider self-hosting or doing these sorts of things that are potentially more technical for the majority of the user bases of you know whether it be
31:01
31 minutes, 1 second
Blue Sky or any emerging platforms on the protocol. Uh so I'm just curious what your general thoughts are on you
31:08
31 minutes, 8 seconds
know how can we make people aware that this is how the network works rather than a Twitter 2.0 centralized service.
31:17
31 minutes, 17 seconds
Yep. Totally. Uh I mean I think that the the thing that you can't do is force users to understand that, right? You have to give them the ability to
31:26
31 minutes, 26 seconds
gradually ramp up. And so it's like you have an account on the protocol, that's great. Okay, you set up a recovery key, even better. You're doing backups of
31:34
31 minutes, 34 seconds
your repo, even better. Okay, you're self-hosting your account. Wow, even better. Okay, you're doing that and you're using a self-hosted uh
31:41
31 minutes, 41 seconds
application. Awesome. Even better. Uh so you need to give this on-ramp for users to gradually adopt more responsibility
31:49
31 minutes, 49 seconds
for their experience in the protocol but you can't thrust all of it on them at the beginning and then it's just a question of uh building good UX around
31:57
31 minutes, 57 seconds
that making it approachable figuring out how to explain the ideas. Um I think that there are ways to do it. It is a hard question. I joke that blue sky is a
32:06
32 minutes, 6 seconds
like hat that looks like Twitter sitting on top of a you know distributed systems company but like it's it's a skeworphism
32:14
32 minutes, 14 seconds
that people can understand and recognize and gradually introduce people to how weird we can get. Yeah.
32:25
32 minutes, 25 seconds
Yeah. Go see Miss Boba as well. Hey Daniel. Uh Sam here from Fedica. I have a question about car files. Uh do you
32:33
32 minutes, 33 seconds
guys plan to continue to use car files uh in the future? And if so uh because they're a little difficult to work with
32:42
32 minutes, 42 seconds
from the developer community and if so do you plan to make it easier just like how Jetream made it simpler than than
32:49
32 minutes, 49 seconds
the fire hose? Yeah, I do think that we plan to use car files going forward. Uh we really want to write our own car
32:56
32 minutes, 56 seconds
library that has better interfaces. The current one is not very great. Uh so uh yes we do want to use them. Yes we want
33:04
33 minutes, 4 seconds
to make them easier to use. And then on your question of Jetream my hope is that we evolve Jetream into a a larger
33:11
33 minutes, 11 seconds
broader purpose like sync facilitation service that not only provides this like JSON API of records but also facilitates
33:18
33 minutes, 18 seconds
backfill and repo sync and some of these headier sync operations.
33:28
33 minutes, 28 seconds
Other questions? Otherwise, I definitely have a question for Daniel.
33:31
33 minutes, 31 seconds
Um, the I'm I'm interested in hearing more about how you define the fear and the sort of threat modeling, right? Like
33:39
33 minutes, 39 seconds
one of the exciting parts about the approach that y'all are taking as a company is something that um Brian
33:46
33 minutes, 46 seconds
Fitzpatrick talked about um with um the data liberation front at Google and the idea that um companies should always
33:54
33 minutes, 54 seconds
make it possible for their users to be able to migrate away because that is the way that you can say like your users are affirmatively committed to the decisions
34:03
34 minutes, 3 seconds
that you're making and why you're you know that that y'all can maintain alignment and that you're not just trapping them there, right? Um, when it
34:12
34 minutes, 12 seconds
comes to how you're thinking about that both at the protocol design level, like how evil are you imagining that your apps can get, like how does that
34:19
34 minutes, 19 seconds
actually tangibly manifest in the way that you're thinking and sort of threat modeling the types of bad actions that you could take? Yeah, totally.
34:28
34 minutes, 28 seconds
Um, I mean, I don't know how evil they can get. There's messed up people that that's kind of thing, right? Like the the the threat modeling part about this is like, you know, I worked with
34:36
34 minutes, 36 seconds
journalists for a really long time and journalists can have any conceivable threat model from like I only work on public records and like I don't care what people see all the way to like if
34:45
34 minutes, 45 seconds
anybody finds my files like somebody's going to get their fingernails torn out, right? Um and and so like the the thing that impresses me about the way that you
34:53
34 minutes, 53 seconds
all think about design is the decomposition of the problem space and and so uh uh figuring out how to teslate out like which problems you want to
35:01
35 minutes, 1 second
threat model I think is an interesting way that you go about this. Yeah, totally. I I think like the big way that I think about it is that
35:09
35 minutes, 9 seconds
uh you since you always have that credible exit, it prevents one rentseeking behaviors where people can abuse their position as like some
35:18
35 minutes, 18 seconds
central provider and uh it also uh gives an incentive not to abuse users on your service and it also gives the ability to
35:27
35 minutes, 27 seconds
people for people to make you know very bespoke services to fit certain needs.
35:31
35 minutes, 31 seconds
So if someone has a very high privacy need, they can use a service like that and interoperate with the rest of the network. So to I don't know, maybe to
35:39
35 minutes, 39 seconds
compare it to like messaging systems or something, it's like uh you know, my aunt is never going to use signal. I there's nothing I could do to convince
35:46
35 minutes, 46 seconds
her to do that. But if I'm very privacy conscious, it would be great if I could use signal and still talk to my aunt. Uh and similarly it's like if someone has
35:55
35 minutes, 55 seconds
like very high privacy needs uh with this choice of service uh or high privacy or data security needs with this ability to choose your service hopefully
36:04
36 minutes, 4 seconds
you could choose a PDS service provider that is really tailored to those.
36:09
36 minutes, 9 seconds
Um so yeah I think that the the the choice of service gives you a lot of optionality on that. Cool. Thanks. All right. Other questions? Otherwise we'll say yeah come on over.
36:24
36 minutes, 24 seconds
Proton PDS Daniel from the Bay. So, I was I had a
36:33
36 minutes, 33 seconds
conversation about this earlier in the week. Um Boris had chimed in, but um more so a non-technical question. Um how
36:40
36 minutes, 40 seconds
has Blue Sky discussed how to signal to people that they can sign in um to like an app for example with their PDS? Like
36:49
36 minutes, 49 seconds
let's say for example, you meet someone who doesn't know what Blue Sky is, but there's a they're signing into an app or they're about to they're at the login screen of an app that is on, you know,
36:58
36 minutes, 58 seconds
the AT Pro protocol. How do you signal to them that they can s that they can sign into um sign into that app because
37:06
37 minutes, 6 seconds
you can't say sign into Blue Sky? Yeah, totally. What is that, right? You know, that's a great question. That's a UX thing we need to figure out. If people
37:13
37 minutes, 13 seconds
have good ideas on this, I definitely want to talk with them about it. But uh the great thing about app proto identity is that it can serve as this universal
37:21
37 minutes, 21 seconds
identifier both for social apps that run on top of the app protocol but also for non uh app protocol apps and uh figuring
37:29
37 minutes, 29 seconds
out some UX pattern to be able to posit it like sign in with atmosphere but that's kind of wordy and confusing. So we got to figure out a good one.
37:37
37 minutes, 37 seconds
Yes, sign in with cloudy ad symbol because in the GitHub discussion, I know there was a specific mention that um I
37:46
37 minutes, 46 seconds
believe it was Jay that um mentioned that she didn't want at Proto actually used as like for branding.
37:56
37 minutes, 56 seconds
Some someone in the someone mentioned it's either someone mentioned that um leader said that um she didn't want at
38:04
38 minutes, 4 seconds
Proto used for branding or that branding um that at Proto shouldn't be used for branding. So I was just like well what
38:12
38 minutes, 12 seconds
what should we do you know? Yeah I'm actually not quite sure on the context of that. So maybe maybe we can chat about it after. But uh we definitely do
38:20
38 minutes, 20 seconds
want people to understand the atmosphere and have a coherent conception of the atmosphere and know that you can log into stuff with the atmosphere. And I do think a lot of that is going to be the brand and apps being able to use that
38:29
38 minutes, 29 seconds
brand. This is also one of the things that came up in one of the tech talks that uh Boris hosted a couple of months ago. And I I definitely also see that this is a community responsibility in
38:38
38 minutes, 38 seconds
part and this is one of those things that we can work we need to be able to work on together. But also you'll have the largest install base, right? So, yes, we're interested in engaging y'all
38:46
38 minutes, 46 seconds
on that more. Um, all right. Uh, I think we've got time for one more question if anybody's got All right.
38:55
38 minutes, 55 seconds
Okay, we can do two.
39:00
39 minutes
Um, hi. Yeah. Uh so I think it's in the lexicon community and maybe Nick can correct me but there is a github discussion on this topic you know how to
39:09
39 minutes, 9 seconds
uh improve the UX and the UI flow of logging through other methods. Uh I forgot the link but I think it should be
39:17
39 minutes, 17 seconds
under one discussions on the lexicon community GitHub repo. Cool. Thanks.
39:25
39 minutes, 25 seconds
Hi. There was a long list of like previous references that informed the atproto ethos. I'm curious if there are protocols out there today that you think
39:34
39 minutes, 34 seconds
are at ProtoE or share similar values or just things like that.
39:39
39 minutes, 39 seconds
[Music]
39:42
39 minutes, 42 seconds
Um, I'd say surprisingly Noir actually has a similar architecture to app proto in terms of having uh relays
39:50
39 minutes, 50 seconds
broadcasting stuff out and aggregating it in applications although they don't really have the uh like very rich
39:58
39 minutes, 58 seconds
backends that uh app proto has with app views. Um, but it sort of has a similar shape of information flow.
40:08
40 minutes, 8 seconds
Um I don't know the web still.
40:14
40 minutes, 14 seconds
Yeah. Um thank you very much. Uh this is great. Uh appreciate all the questions as well. Sweet. Thanks guys.