This first blog is about this website and how it was made.
When I first began looking into web development it was for a friend,
they requested a small website to do time tracking for studying.
I used NodeJS and Express to build a super simple website, and everything was
fine. They never ended up using that website because I wasn't aware of
Javascript canvas at the time to write proper data visualizations, and I then
got too busy with school.
However, as I got closer to graduating I realized that it would be nice to have
a personal website. Furthermore I recently took a class where we used javascript
and webgl to do computer graphics, so I felt fairly confident at building
more dynamic and interesting web pages.
Since I had more time and the ability to set my own pace I decided to use
C to write my own custom webserver. I am a huge fan of the C programming language,
and I already had my own mini library of helpers to make C more convenient.
So the first version of this webserver was essentially what you'd find in a
undergrad Systems Programming class.
Essentially it just waited until it got a connection, read the message into a buffer,
and then parsed the http, producing a map of headers and values, as well as a
URI in the form of a list of strings.
Then I would read the uri, open the requested static file and send it over the network
before closing the connection. So at this point the code looked something like this.
int socket = socket();
listen(socket);
Buffer data = {};
while (TRUE) {
int fd = accept(socket);
while (!data.size || data.back = "\r\n\r\n") {
Resize(data);
read(fd, data);
}
HttpReq r = parse(data);
SendRes(r);
}
Making The Server Async
The first major issue with this approach is that the server throughput is pathetically
low. Single threaded blocking IO means that I spend most of my time waiting for the OS
to fill buffers and little time actually handling requests. So I needed to go asynchronous,
in order to reduce wait times.
int socket = socket();
listen(socket);
FDList list;
list.add(socket);
while (TRUE) {
int num_events = poll(list);
for (int i = 0; i < num_events; i++) {
handle_event(list[i]);
}
}
Psuedo Code for Poll based Server
There were a few obvious options, I had previously written a webserver using Linux Poll,
and I was aware that EPoll was an option. If I was a little more adventurous there was
also Linux aio interface. This was better as it actually exposed the real asynchronous framework,
but it still wasn't that interesting to me.
I had heard of io_uring before, and the architecture of an application using io_uring seemed
more interesting and less complicated than a callback based setup. I decided to go with
liburing and io_uring since it seemed the most interesting and flexible choice, although
I do admit I didn't look particularly deeply into Linux aio.
The Architecture
I went through several versions of the architecture as I learned more about
io_uring and as I encountered various problems.
The first version, and all other versions, utilzed an event system. Essentially
I thought the best method was to wait until any io operation finished, then
iterate through all io operations performing the necessary actions, before submitting
all the io at once. That way all io ops go through roughly the same steps and don't
require a ton of book-keeping.
My first attempt at this was essentially a loop with a switch statement inside. For each
request I remembered the file descriptor used, a buffer for the incoming request, and
a buffer for the outgoing response. These buffers were dynamic. I could have gone with static
buffers, which in some ways would have been ideal, but I assumed the buffers would grow to some
maximum size, and I didn't want to have to constantly adjust the sizes of the buffers.
This attempt was fairly stable, although it wasn't particularly robust. And when I wanted to
begin adding blogs I realized that it wasn't expandable in the ways I required.
Attempt 2
This version built directly off of the previous. I decided to pack as much of the non
buisness logic code into an iterator. The idea being that the main function would contain mostly
business logic and that the connection management, IO, and such would reside in what was essentially
a static library.
This worked fairly well, I completely put Accept events, Close events, timeouts, etc. into a
single function. From here I planned out a templating system. I could have gone with a more
robust and flexible standard like html modules, or even React. In the end I settled on something
with less power delibrately, I thought that I would prefer most of the logic be in the server itself,
rather than built on top of something which is less optimized and more complicated.
This version was fairly good, but as I added timeouts and started stress testing cracks began to show.
Sometimes requests would be dropped with no easily debuggable cause. I sometimes also ran into spontaneous
crashes, and I even found multiple race conditions where connections could be erroneously closed.
Attempt 3
Keeping in mind all the lessons I had learned I decided to see if io_uring had any better
options for me. While researching I found two things which would massively simplify the implementation,
thus reducing bugs.
The first was direct file tables. io_uring, has an option to register a number of file descriptors in a table.
Doing this meant that io operations could skip the fdget and fdput operations increasing performance. However,
for me there was a far more important benefit, the operating system would track which file descriptors
in the table were active and which were free, meaning I could strip out all the logic
to allocate resources to connections which removed an entire class of bugs.
The second was the discovery of pre-registered buffers. Using these buffers you could
make the operating system aware of buffers to be used in recv, read, or any operation which
writes to userspace memory. Using this I could replace the dynamic buffers for storing requests
with easier to manage fixed size buffers, so long as I could parse http lazily. To do that
I found the RFC with the http spec and created a simple state machine. This meant I could
utilize fixed size buffers and could avoid having to store the entire request.
This was a massive benefit as well. In practice almost no requests are more than 4kb, but the fact
that it is no longer necessary to accomadate the edge case of needing more memory meant that I could remove
all the logic and allocations required to handle dynamicly sized requests.
In addition to these technical discoveries, while I was rewriting the core event loop I came to a realization.
I don't need to support multiple event types being surfaced to the user, rather there is a single event type
an http message. Once I recognized this I could remove the switch statement in the business logic and just
have the single block for handling messages. This also simplified the event loop itself since now the business
logic and async management can be completely seperate.
Conclusion
I learned quite a few new things while working on this website. I learned the basics
of io_uring, as well as explored more the possibility space when it comes to
async programming.
There were also some interesting design decisions I decided to gloss over in this first post,
including the Http Parser, how I handle files with 0 per request syscalls, and my personal
C utility libraries including a custom build system.