Monday, October 4, 2010

Writing your own storage engine for Memcached

I am working full time on membase, which utilize the "engine interface" we're adding to Memcached. Being the one who designed the API and wrote the documentation, I can say that we do need more (and better) documentation without insulting anyone. This blog entry will be the first entry in mini-tutorial on how to write your own storage engine. I will try to cover all aspects of the engine interface while we're building an engine that stores all of the keys on files on the server.

This entry will cover the basic steps of setting up your development environment and cover the lifecycle of the engine.

Set up the development environment

The easiest way to get "up'n'running" is to install my development branch of the engine interface. Just execute the following commands:

$ git clone git://github.com/trondn/memcached.git
$ cd memcached
$ git -b engine origin/engine
$ ./config/autorun.sh
$ ./configure --prefix=/opt/memcached
$ make all install
     

Lets verify that the server works by executing the following commands:

$ /opt/memcached/bin/memcached -E default_engine.so &
$ echo version | nc localhost 11211
VERSION 1.3.3_433_g82fb476     ≶-- you may get another output string....
$ fg
$ ctrl-C
     

Creating the filesystem engine

You might want to use autoconf to build your engine, but setting up autoconf is way beyond the scope of this tutorial. Let's just use the following Makefile instead.

ROOT=/opt/memcached
INCLUDE=-I${ROOT}/include

#CC = gcc
#CFLAGS=-std=gnu99 -g -DNDEBUG -fno-strict-aliasing -Wall \
# -Wstrict-prototypes -Wmissing-prototypes -Wmissing-declarations \
# -Wredundant-decls \
# ${INCLUDE} -DHAVE_CONFIG_H
#LDFLAGS=-shared

CC=cc
CFLAGS=-I${ROOT}/include -m64 -xldscope=hidden -mt -g \
      -errfmt=error -errwarn -errshort=tags  -KPIC
LDFLAGS=-G -z defs -m64 -mt

all: .libs/fs_engine.so

install: all
 ${CP} .libs/fs_engine.so ${ROOT}/lib

SRC = fs_engine.c
OBJS = ${SRC:%.c=.libs/%.o}

.libs/fs_engine.so: .libs $(OBJS)
 ${LINK.c} -o $@ ${OBJS}

.libs:; -@mkdir $@

.libs/%.o: %.c
 ${COMPILE.c} $< -o $@   clean:  $(RM) .libs/fs_engine.so $(OBJS)       

I am doing most of my development on Solaris using the Sun Studio compilers, but I have added a section with settings for gcc there if you're using gcc. Just comment out lines for CC, CFLAGS and LDFLAGS and remove the # for the gcc alternatives.

In order for memcached to utilize your storage engine it needs to first load your module, and then create an instance the engine. You use the -E option to memcached to specify the name of the module memcached should load. With the module loaded memcached will look for a symbol named create_instance in the module to create an handle memcached can use to communicate with the engine. This is the first function we need to create, and it should have the following signature:

MEMCACHED_PUBLIC_API
ENGINE_ERROR_CODE create_instance(uint64_t interface, GET_SERVER_API get_server_api, ENGINE_HANDLE **handle);
     

The purpose of this function is to provide the server a handle to our module, but we should not perform any kind of initialization of our engine yet. The reason for that is because the memcached server may not support the version of the API we provide. The intention is that the server should notify the engine with the "highest" interface version it supports through interface, and the engine must return a descriptor to one of those interfaces through the handle. If the engine don't support any of those interfaces it should return ENGINE_ENOTSUP.

So let's go ahead and define a engine descriptor for our example engine and create an implementation for create_instance:

struct fs_engine {
  ENGINE_HANDLE_V1 engine;
  /* We're going to extend this structure later on */
};

MEMCACHED_PUBLIC_API
ENGINE_ERROR_CODE create_instance(uint64_t interface,
                                 GET_SERVER_API get_server_api,
                                 ENGINE_HANDLE **handle) {
  /*
   * Verify that the interface from the server is one we support. Right now
   * there is only one interface, so we would accept all of them (and it would
   * be up to the server to refuse us... I'm adding the test here so you
   * get the picture..
   */
  if (interface == 0) {
     return ENGINE_ENOTSUP;
  }

  /*
   * Allocate memory for the engine descriptor. I'm no big fan of using
   * global variables, because that might create problems later on if
   * we later on decide to create multiple instances of the same engine.
   * Better to be on the safe side from day one...
   */
  struct fs_engine *h = calloc(1, sizeof(*h));
  if (h == NULL) {
     return ENGINE_ENOMEM;
  }

  /*
   * We're going to implement the first version of the engine API, so
   * we need to inform the memcached core what kind of structure it should
   * expect
   */
  h->engine.interface.interface = 1;

  /*
   * Map the API entry points to our functions that implement them.
   */
  h->engine.initialize = fs_initialize;
  h->engine.destroy = fs_destroy;

  /* Pass the handle back to the core */
  *handle = (ENGINE_HANDLE*)h;

  return ENGINE_SUCCESS;
}
     

If the interface we provide in create_instance is dropped from the supported interfaces in memcached, the core will call destroy() immediately. The memcached core guarantees that it will never use any pointers returned from the engine when destroy() is called.

So let's go ahead and implement our destroy() function. If you look at our implementation of create_instance you will see that we mapped destroy() to a function named fs_destroy():

static void fs_destroy(ENGINE_HANDLE* handle) {
  /* Release the memory allocated for the engine descriptor */
  free(handle);
}
     

If the core implements the interface we specify, the core will call a the initialize() method. This is the time where you should do all sort of initialization in your engine (like connecting to a database, initializing mutexes etc). The initialize function is called only once per instance returned from create_instance (even if the memcached core use multiple threads). The core will not call any other functions in the api before the initialization method returns.

We don't need any kind of initialization at this moment, so we can use the following initialization code:

static ENGINE_ERROR_CODE fs_initialize(ENGINE_HANDLE* handle,
                                      const char* config_str) {
  return ENGINE_SUCCESS;
}
     

If the engine returns anything else than ENGINE_SUCCESS, the memcached core will refuse to use the engine and call destroy()

In the next blog entry we will start adding functionality so that we can load our engine and handle commands from the client.

Tuesday, September 7, 2010

Birkebeinerrittet

I got inspired to attend the worlds largest cross country bicycle race named Birkebeinerrittet when I saw a TV show from the race a couple of years back. Everyone that knows me knows that I'm probably as far as you can get from a top athlete (I'm a hacker ;-), so I borrowed a bike from a good friend of mine earlier this summer and started to prepare for the 96km bike ride over the mountain from Rena to Lillehammer.

I had a nice trip over the mountain, but unfortunately it rained most of the time up there so I didn't get to see much of the nature. On the bright side I was one of the lucky ones that chose to attend the race on Friday instead of the main event on Saturday since it kept on raining that night and the next day causing the track to be even more muddy (and the temperature kept on dropping).

Because this was the first time I attended the race I didn't know what to expect, so I was a bit afraid if I would run out of energy. After the race I felt I had more to give, so I am really motivated for attending next year as well. There is a separate quota for foreigners who want to attend the race, so please join me next year!

Wednesday, July 28, 2010

libmemcached on win32

It's a fact that people love their development platform and want to stick with it. I'm a die hard Solaris fan, and would never dream of switching to something else. I've heard that there is a crowd out there that likes to work on other systems like Windows, MacOSX, BSD and Linux. That the developer use a platform during development doesn't necessary mean that the target product will run on the platform, but the developer is more productive on that platform.

People can argue as much as they want, but there is a large crowd of developers using Windows. There is also a large number of systems running some version of Windows, so enabling them to use the projects I'm working on is a good thing. Earlier today I pushed a branch that adds support for building libmemcached into a dll on Windows.

90% of the source code in libmemcached is just "logic" that applies for all platforms, but there is a small part of the code that interacts tightly with the operating system. "Everything" is a file descriptor on Unix systems, but Windows got their own subsystem for sockets called WinSock. In order to avoid getting tonns of #ifdefs all over the code, I defined memcached_socket_t to represent a socket object, and used "the WinSock way" to implement the code. It is pretty easy to map the WinSock code to work on Unix systems with a couple of macros.

Well, enough talk. If you're interested in the details you can check out the branch on Launchpad.

The easiest way for you to test out the code is to install the fullinstall of msysgit. Unfortunately it doesn't come with all of the tools needed to build libmemcached (you can't generate a configure script and generate the documentation). This means that you cannot build the development branch unless you got another machine where you can generate the configure script. I am exporting parts of my ZFS filesystem via CIFS, so I generated the configure script on Solaris.

Building libmemcached with mingw is just as easy as on your favorite platform:

$ ./configure --with-memcached= --without-docs
$ make all install

I haven't fixed the test suite yet, so you have to wait a bit longer before you can run make test ;)

Tuesday, March 30, 2010

Building Memcached on Windows

I like to be able to compile my software on multiple platforms with different compilers, because it force me to write that complies to standards (if not you'll most likely have a spaghetti of #ifdefs all over your code). One of the bonuses you get from using multiple compilers is that you'll get "more eyes" looking at your code, and they may warn you on different things. Supporting multiple platforms will also make it easier to ensure that you don't have 'hidden bombs' in your code (byte order, alignment, assumptions that you can dereference a NIL pointer etc)...

Building software that runs on different flavors of Unix/Linux isn't that hard, but adding support for Microsoft Windows offers some challenges. Among the challenges are:

  1. Microsoft Windows doesn't provide the headerfiles found on most Unix-like systems (unistd.h etc).
  2. Sockets is implemented in a separate "subsystem" on windows.
  3. Win32 provides another threads/mutex implementation than what most software written for Unix-like systems use (pthreads).
  4. Microsoft doesn't supply a C99 compiler. One would think that it would be safe to assume that one could use a C standard that is more than 10 years old....

I am really glad to announce that we just pushed a number of changesets to memcached to allow you to build memcached on windows "just as easy" as you would on your favorite Unix/Linux system. This means that as of today users using Microsoft Windows is no longer stuck with an ancient version!

`The first thing you need to do in order to build memcached on Windows is to install the "fullversion" of msysgit. In addition to installing git, it also installs a compiler capable of building a 32 bit version of memcached.

So let's go ahead and build the dependencies we need to get our 32bit memcached version up and running!

The first thing up is libevent. You should download the latest 2.x release available (there is a bug in all versions up to (and including) 2.0.4, so unless there is a new one out there you can grab a development version I pushed to libevent-2.0.4-alpha-dev.tar.gz), and install it with the following commands (from within your msysgit-shell). Please note that I'm using /tmp because I've had problems using my "home directory" because of spaces in the directory name (or at least I guess that's the reason ;-) ):

$ cd /tmp
$ tar xfz libevent-2.0.4-alpha-dev.tar.gz
$ cd libevent-2.0.4-alpha-dev
$ ./configure --prefix=/usr/local
$ make all                    (this will fail with an error, but don't care about that.. it's is in the example code)
$ make install             (this will fail with an error, but don't care about that.. it's just an example)

It's time to start build memcached!!! Let's check out the source code and build it!

$ git clone git://github.com/trondn/memcached.git
$ git checkout -t origin/engine
$ make -f win32/Makefile.mingw

You should now be able to start memcached with the following command:

$ ./memcached.exe -E ./libs/default_engine.so

Go ahead and telnet to port 11211 and issue the "stats" command to verify that it works!

You don't need to install msysgit on the systems where you want to run memcached, but you do need to include pthreadGC2.dll from the msysgit distribution. That means that if you want to run memcached on another machine you need to copy the following files: memcached.exe .libs/default_engine.so and /mingw/bin/pthreadGC2.dll (Just place them in the same directory on the machine you want to run memcached on :-)

But wait, the world is 64 bit by now. Lets build a 64 bit binary instead

The msysgit we installed previously isn't capable of building a 64-bit binary, so we need to install a new compiler. Download the a bundle from http://sourceforge.net/projects/mingw-w64/files/ (I used mingw-w64-bin_i686-mingw_20100129.zip ). You can "install" it by running the following commands in our msysgit shell:

$ mkdir /mingw64
$ cd /mingw64
$ unzip /mingw-w64-bin_i686-minw_20100129.zip
$ export PATH=/mingw64/bin:/usr/local/bin:$PATH

So let's start compiling libevent:

$ cd /tmp
$ tar xfz libevent-2.0.4-alpha-dev.tar.gz
$ cd libevent-2.0.4-alpha-dev
$ ./configure --prefix=/usr/local --host=x86_64-w64-mingw32 --build=i686-pc-mingw32
$ make all
$ make install

The compiler we just downloaded didn't come with a 64-bit version of pthreads, so we have to download and build that ourself. I've pushed a package to my machine at: pthreads-w64-2-8-0-release.tar.gz

$ cd /tmp
$ tar xfz pthreads-w64-2-8-0-release.tar.gz
$ cd pthreads-w64-2-8-0-release
$ make clean GC CC=x86_64-w64-mingw32-gcc
$ cp pthread.h semaphore.h sched.h /usr/local/include
$ cp libpthreadGC2.a /usr/local/lib/libpthread.a
$ cp pthreadGC2.dll /usr/local/lib
$ cp pthreadGC2.dll /usr/local/bin

And finally memcached with:

$ git clone git://github.com/trondn/memcached.git
$ git checkout -t origin/engine
$ make -f win32/Makefile.mingw CC=x86_64-w64-mingw32-gcc

So how does it perform?

I'm pretty sure a lot of the "Linux fanboys" are ready to jump in and tell how much faster the Linux version is. I have to admit that I was a bit skeptical in the beginning if the layers we added to get "Unix compatibility" had any performance impact. The only way to know for sure is to run a benchmark to measure the performance.

I ran a small benchmark where I had my client create 512 connections to the server, and then randomly choosing a connection and perform a get or a set operation on one of the 100 000 items I had stored in the server. I ran 500 000 operation (33% of the operations are set operations), and calculated the average time. I can dual-boot one of my machines into Windows 7 and Red Hat Enterprise Linux, so my results shows the "out of the box"-numbers between Windows 7 and Red Hat Enterprise Linux running on the same hardware. I used my OpenSolaris box to drive the test (connected to the same switch). Posting numbers is always "dangerous", because people put too much into the absolute numbers. The intention with my graphs is to show that the version running on Microsoft Windows is pretty much "on par" with the Linux version.

256 byte userdata

512 byte userdata

1024 byte userdata

Wednesday, March 3, 2010

Memcached with SASL on OpenSolaris - part 2

It turns out that some systems doesn't support shadow.h and fgetspent, so I just updated the patch to no longer require them.. To create the file, simply do: echo "myuser:mypass" >> my_sasl_pwdb Happy hacking

Thursday, February 25, 2010

Memcached with SASL on OpenSolaris

You may have tried to build memcached with SASL support on OpenSolaris with the following result:

trond@opensolaris< ./configure --enable-sasl
[ ... cut ... ]
checking sasl/sasl.h usability... yes
checking sasl/sasl.h presence... yes
checking for sasl/sasl.h... yes
checking for library containing sasl_server_init... no
configure: error: Failed to locate the library containing sasl_server_init

This is because configure only tries to look for sasl_server_init in libsasl2, and OpenSolaris use libsasl instead. Yesterday I pushed a fix to my github repository that search for the symbol in libsasl as well.

Configuring SASL may also be a challenge (I want to spend my time writing code, not be a system administrator), so I decided to add support for plaintext passwords as well. You probably don't want to use this on your production servers, but it comes in really handy if you just want to test your favorite client.

You enable support for plaintext password by passing --enable-sasl-pwdb to configure. I didn't want to spend any time to writing a new parser or come up with a new file format, so I decided to use fgetspent_r to read the password file. This means that as long as you follow the format for a shadow file, you're good to go :-) You have to set the name of the file to use as the password file in the environment variable MEMCACHED_SASL_PWDB:

trond@opensolaris> echo "myname:mypass:::::::" > /tmp/memcached-sasl-db
trond@opensolaris> export MEMCACHED_SASL_PWDB=/tmp/memcached-sasl-db

With a password file in place, you have to create a config file for SASL to instruct it to use plain text password authentication:

trond@opensolaris> echo "mech_list: plain" > memcached.conf

If you don't want to install this as the global configuration for memcached, you should specify the location of the file in SASL_CONF_PATH:

trond@opensolaris> export SASL_CONF_PATH=`pwd`/memcached.conf

You then start the memcached deamon with "-S" to enable SASL authentication:

trond@opensolaris> ./memcached -S -d

So let's run some commands to the server and see how this works. I'm using the SASL support I implemented in libmemcached (not integrated yet, but you may download it from https://code.launchpad.net/~trond-norbye/libmemcached/sasl):

trond@opensolaris> ./memcp --servers=localhost:11211 --binary \
                                                    --username=myname --password=inncorrect \
                                                    memcp.c
memcp: memcp.c: memcache error AUTHENTICATION FAILURE
trond@opensolaris> ./memcp --servers=localhost:11211 --binary \
                                                    --username=myname --password=mypass \
                                                    memcp.c
trond@opensolaris> ./memcat --servers=localhost:11211 --binary \
                                                     --username=myname --password=mypass \
                                                     memcp.c
[ ... output of memcp.c ... ]
That's all for now. Happy hacking :-)

Thursday, February 11, 2010

New opportunities

I've been really quiet lately, so I guess I should come with a short update on what I'm up to. As of 2010 I am no longer a Sun Microsystems employee, so I spent January in Mountain View in the offices of my new employer: NorthScale. I'm currently working on getting up to speed on my new tasks, so I'll be back with more information later.