First Blog!

This first blog is about this website and how it was made. When I first began looking into web development it was for a friend, they requested a small website to do time tracking for studying. I used NodeJS and Express to build a super simple website, and everything was fine. They never ended up using that website because I wasn't aware of Javascript canvas at the time to write proper data visualizations, and I then got too busy with school. However, as I got closer to graduating I realized that it would be nice to have a personal website. Furthermore I recently took a class where we used javascript and webgl to do computer graphics, so I felt fairly confident at building more dynamic and interesting web pages. Since I had more time and the ability to set my own pace I decided to use C to write my own custom webserver. I am a huge fan of the C programming language, and I already had my own mini library of helpers to make C more convenient. So the first version of this webserver was essentially what you'd find in a undergrad Systems Programming class.
Essentially it just waited until it got a connection, read the message into a buffer, and then parsed the http, producing a map of headers and values, as well as a URI in the form of a list of strings. Then I would read the uri, open the requested static file and send it over the network before closing the connection. So at this point the code looked something like this.

    int socket = socket();
    listen(socket);

    Buffer data = {};
    
    while (TRUE) {
        int fd = accept(socket); 
         
        while (!data.size || data.back = "\r\n\r\n") {
            Resize(data);
            read(fd, data);
        }

        HttpReq r = parse(data);
        SendRes(r);
    }

Making The Server Async

The first major issue with this approach is that the server throughput is pathetically low. Single threaded blocking IO means that I spend most of my time waiting for the OS to fill buffers and little time actually handling requests. So I needed to go asynchronous, in order to reduce wait times.


    int socket = socket();
    listen(socket);

    FDList list;
    list.add(socket);

    while (TRUE) {
        int num_events = poll(list);

        for (int i = 0; i < num_events; i++) {
            handle_event(list[i]);
        }
    }


Psuedo Code for Poll based Server
There were a few obvious options, I had previously written a webserver using Linux Poll, and I was aware that EPoll was an option. If I was a little more adventurous there was also Linux aio interface. This was better as it actually exposed the real asynchronous framework, but it still wasn't that interesting to me. I had heard of io_uring before, and the architecture of an application using io_uring seemed more interesting and less complicated than a callback based setup. I decided to go with liburing and io_uring since it seemed the most interesting and flexible choice, although I do admit I didn't look particularly deeply into Linux aio.

The Architecture

I went through several versions of the architecture as I learned more about io_uring and as I encountered various problems. The first version, and all other versions, utilzed an event system. Essentially I thought the best method was to wait until any io operation finished, then iterate through all io operations performing the necessary actions, before submitting all the io at once. That way all io ops go through roughly the same steps and don't require a ton of book-keeping. My first attempt at this was essentially a loop with a switch statement inside. For each request I remembered the file descriptor used, a buffer for the incoming request, and a buffer for the outgoing response. These buffers were dynamic. I could have gone with static buffers, which in some ways would have been ideal, but I assumed the buffers would grow to some maximum size, and I didn't want to have to constantly adjust the sizes of the buffers. This attempt was fairly stable, although it wasn't particularly robust. And when I wanted to begin adding blogs I realized that it wasn't expandable in the ways I required.

Attempt 2

This version built directly off of the previous. I decided to pack as much of the non buisness logic code into an iterator. The idea being that the main function would contain mostly business logic and that the connection management, IO, and such would reside in what was essentially a static library. This worked fairly well, I completely put Accept events, Close events, timeouts, etc. into a single function. From here I planned out a templating system. I could have gone with a more robust and flexible standard like html modules, or even React. In the end I settled on something with less power delibrately, I thought that I would prefer most of the logic be in the server itself, rather than built on top of something which is less optimized and more complicated. This version was fairly good, but as I added timeouts and started stress testing cracks began to show. Sometimes requests would be dropped with no easily debuggable cause. I sometimes also ran into spontaneous crashes, and I even found multiple race conditions where connections could be erroneously closed.

Attempt 3

Keeping in mind all the lessons I had learned I decided to see if io_uring had any better options for me. While researching I found two things which would massively simplify the implementation, thus reducing bugs. The first was direct file tables. io_uring, has an option to register a number of file descriptors in a table. Doing this meant that io operations could skip the fdget and fdput operations increasing performance. However, for me there was a far more important benefit, the operating system would track which file descriptors in the table were active and which were free, meaning I could strip out all the logic to allocate resources to connections which removed an entire class of bugs. The second was the discovery of pre-registered buffers. Using these buffers you could make the operating system aware of buffers to be used in recv, read, or any operation which writes to userspace memory. Using this I could replace the dynamic buffers for storing requests with easier to manage fixed size buffers, so long as I could parse http lazily. To do that I found the RFC with the http spec and created a simple state machine. This meant I could utilize fixed size buffers and could avoid having to store the entire request. This was a massive benefit as well. In practice almost no requests are more than 4kb, but the fact that it is no longer necessary to accomadate the edge case of needing more memory meant that I could remove all the logic and allocations required to handle dynamicly sized requests. In addition to these technical discoveries, while I was rewriting the core event loop I came to a realization. I don't need to support multiple event types being surfaced to the user, rather there is a single event type an http message. Once I recognized this I could remove the switch statement in the business logic and just have the single block for handling messages. This also simplified the event loop itself since now the business logic and async management can be completely seperate.

Conclusion

I learned quite a few new things while working on this website. I learned the basics of io_uring, as well as explored more the possibility space when it comes to async programming. There were also some interesting design decisions I decided to gloss over in this first post, including the Http Parser, how I handle files with 0 per request syscalls, and my personal C utility libraries including a custom build system.