Archive for the ‘Cube Life’ Category

Swatting Flies is an Annoying Thing to Do

Tuesday, August 17th, 2010

cubeLifeView.gif

I've been working (still) on getting more exchange feed codecs into the system, and while it's not really hard work, it takes a little thought, and a lot of attention to detail. So when I get some kibitzing from those that would love to see me fail, but are too afraid to really stand up to this project, it's like swatting flies - not hard, they aren't going to do me any harm, but it's annoying nonetheless.

When it gets bad, I just get up, take a little walk, get a pop, and clear my head. That usually does it. Oh... and getting another feeder done in less than a day makes me feel good. It shows the "flies" that they really might want to take notice of the different way I've put this together. But that's really hoping for too much, I suppose.

Time to get some bug spray.

Once Again – Amazing Progress with a Good Design

Monday, August 16th, 2010

MarketData.jpg

Today I spent all day working on getting two exchange feeders written and tested. This kind of speed is not because I can copy/paste very fast, it's because I've got a solid design that allows me to leverage the work I've already done and customize it very quickly and easily. Given that the last developers of these feeds took months to achieve what I've done in a day, there's a lot to be said about the power of the design. The previous one was particularly ill-suited to this task.

So it was a hard day, but I'm getting a lot closer to the point that I'm caught up with all the exchange feeds we have. At that point, I can look to data enrichment, and really start to add value to the data feeds.

Fantastic Point About the Consequences of Who You Hire

Friday, August 13th, 2010

I was reading this article by Paul Graham about why Yahoo didn't last, and came across this point: (which, by the way, wasn't John Gruber's highlighted section)

In technology, once you have bad programmers, you're doomed. I can't think of an instance where a company has sunk into technical mediocrity and recovered. Good programmers want to work with other good programmers. So once the quality of programmers at your company starts to drop, you enter a death spiral from which there is no recovery.

Paul goes on to talk about the difference in the culture of Google at 500 people and Yahoo at the same point. It's interesting, but not necessarily surprising for someone that's been in this business as long as Paul, and I, have. It's classic: There's no Free Lunch.

There's no way around hiring the best talent you can get. No one thinks getting the cheapest artist is a good idea. Nor the cheapest surgeon. Everyone seems to understand that when you're dealing with an artistic, or especially challenging area of study, and there's one and only one person at the task, that it's a good idea - no, the right idea, to get the best you can. What they think is that coding, like building automobiles, or making frozen pizzas, is something that you can get better at by throwing more people at it.

There are tons of books on this. Even more Harvard Business Studies. It's a deceptively simple lie - programmers are like ants - you just need more. Yup... you keep thinking that. It's a lie, plain and simple.

Everything we humans do has some sense of skill and quality. If you want to be good at something, you have to practice. And not a little. You have to want it. These are the qualities of a good worker - not just that he knows a language, and takes orders. That's a given. You need more.

Sadly, I have a feeling this is never going to be really understood by most people.

The Amazing Power of Really Good Design – And Hard Work

Thursday, August 12th, 2010

Today I was very pleased to see that I could add a second exchange feed to the codebase. Yeah... just one day. Pretty amazing. I know it's primarily due to a good design because the number of lines of code I had to write was very few - on the order of 600 lines, but there's still a little bit of good old hard work to attribute to it as well.

But really, it was the design. What a great design. This is something I'm going to enjoy over and over again as I keep working with this codebase. I need to add in at least six more feeds, but if they are only a day or two per feed, I'm still done long before I had expected to be. Amazing.

So after I had it all done, I looked at the code and realized that when I was "unpacking" the time data from the exchange into milliseconds since epoch, I was making a few system calls, and that was going to come back to bite me later as the loads got higher and higher. The original code looked like:

  /*
   * This method takes the exchange-specific time format and converts it
   * into a timestamp - msec since epoch. This is necessary to parse the
   * timestamp out of the exchange messages as the formats are different.
   */
  uint64_t unpackTime( const char *aCode, uint32_t aSize )
  {
    /*
     * The GIDS format of time is w.r.t. midnight, and a simple, 9-byte
     * field: HHMMSSCCC - so we can parse out this time, but need to add
     * in the offset of the date if we want it w.r.t. epoch.
     */
    uint64_t      timestamp = 0;
 
    // check that we have everything we need
    if ((aCode == NULL) || (aSize < 9)) {
      cLog.warn("[unpackTime] the passed in data was NULL or insufficient "
                "length to do the job. Check on it.");
    } else {
      // first, get the current date/time...
      time_t    when_t = time(NULL);
      struct tm when;
      localtime_r(&when_t, &when);
      // now let's overwrite the hour, min, and sec from the data
      when.tm_hour = (aCode[0] - '0')*10 + (aCode[1] - '0');
      when.tm_min = (aCode[2] - '0')*10 + (aCode[3] - '0');
      when.tm_sec = (aCode[4] - '0')*10 + (aCode[5] - '0');
      // ...and yank the msec while we're at it...
      time_t  msec = ((aCode[6] - '0')*10 + (aCode[7] - '0'))*10 + (aCode[8] - '0');
 
      // now make the msec since epoch from the broken out time
      timestamp = mktime(&when) + msec;
      if (timestamp < 0) {
        // keep it to epoch - that's bad enough
        timestamp = 0;
        // ...and the log the error
        cLog.warn("[unpackTime] unable to create the time based on the "
                  "provided data");
      }
    }
 
    return timestamp;
  }

The problem is that there are two rather costly calls - localtime_r and mktime. They are very necessary, as the ability to calculate milliseconds since epoch is a non-trivial problem, but still... it'd be nice to not have to do that.

So I created two methods: the first was just a rename of this guy:

  /*
   * This method takes the exchange-specific time format and converts it
   * into a timestamp - msec since epoch. This is necessary to parse the
   * timestamp out of the exchange messages as the formats are different.
   */
  uint64_t unpackTimeFromEpoch( const char *aCode, uint32_t aSize )
  {
    // ...
  }

and the second was a much more efficient calculation of the milliseconds since midnight:

  /*
   * This method takes the exchange-specific time format and converts it
   * into a timestamp - msec since midnight. This is necessary to parse
   * the timestamp out of the exchange messages as the formats are
   * different.
   */
  uint64_t unpackTimeFromMidnight( const char *aCode, uint32_t aSize )
  {
    /*
     * The GIDS format of time is w.r.t. midnight, and a simple, 9-byte
     * field: HHMMSSCCC - so we can parse out this time.
     */
    uint64_t      timestamp = 0;
 
    // check that we have everything we need
    if ((aCode == NULL) || (aSize < 9)) {
      cLog.warn("[unpackTimeFromMidnight] the passed in data was NULL "
                "or insufficient length to do the job. Check on it.");
    } else {
      // now let's overwrite the hour, min, and sec from the data
      time_t  hour = (aCode[0] - '0')*10 + (aCode[1] - '0');
      time_t  min = (aCode[2] - '0')*10 + (aCode[3] - '0');
      time_t  sec = (aCode[4] - '0')*10 + (aCode[5] - '0');
      time_t  msec = ((aCode[6] - '0')*10 + (aCode[7] - '0'))*10 + (aCode[8] - '0');
      timestamp = ((hour*60 + min)*60 + sec)*1000 + msec;
      if (timestamp < 0) {
        // keep it to midnight - that's bad enough
        timestamp = 0;
        // ...and the log the error
        cLog.warn("[unpackTimeFromMidnight] unable to create the time "
                  "based on the provided data");
      }
    }
 
    return timestamp;
  }

At this point, I have something that has no system calls in it, and since I'm parsing all these exchange messages, that's going to really pay off in the end. I'm not going to have to do any nasty context switching for these calls - just simple multiplications and additions. I like being able to take the time to go back and clean this kind of stuff up. Makes me feel a lot better about the potential performance issues.

Oh... I forgot... in the rest of my code, I handled the difference in these two by looking at the magnitude of the value. Anything less than "a day" had to be "since midnight" - the rest are "since epoch". Pretty simple.

It works wonderfully!

Over the First Big Hurdle – Really Nice Feeling

Wednesday, August 11th, 2010

MarketData.jpg

This afternoon I can sit back for a minute and look at what I've been doing for the past several weeks as it's gotten to the point that it's tested against data from the exchange, and it's passed those tests. It's only one of about a dozen feeds that I need to handle, but it's the first, and that means that all the infrastructure work I've done - the boost asio sockets... the serialization... the unpacking of exchange data stream... all that is done. Now it's time to put the second codec in the system and see how well it maps to the system I created for one. I'm not really convinced that the design I have now will withstand all the other sources, unmodified, but it's a really good start, and I think it's close.

So I've run this sprint to get the first data feed done, and it's done, and now I find myself exceptionally tired. No surprise there... just a matter of when not if. I've been running on this for a long time without even the slightest break, but it's done now, and I can rest for a minute and then hit it again.

Well... there's my minute's rest... time to get back at it.

Initial Exchange Tests are Wonderfully Successful

Tuesday, August 10th, 2010

MarketData.jpg

This afternoon I was able to finally complete the initial exchange data feed tests on my new ticker plant infrastructure. The data was from a test file the exchange provides, but sent to my UDP receiver via a little test app that we have to simulate the exchange sending the data. I was not surprised to see a few mistakes in the logic - most notably in the parsing, but that's to be expected. Once I got those pointed out, the mistakes were obvious, and the fixes were trivial.

Then it just worked. Fantastic!

The speed was nice, but when you're testing on a single box, it's not really a fair test. I'll have to wait for the two test boxes that are supposed to be provided to me, and then I'll be able to make a much more real-world test. That will be interesting.

What I need to do next is to consolidate the codebase a little - the exchanges send the same data (so they say) on two independent UDP multicast channels, and we've created a dual-headed UDP receiver. We also have a single-headed one. There's too much that's similar, and so I want to consolidate them into a multi-headed UDP receiver and have it just work for from one to, say ten, inputs. The upper limit is really arbitrary, but I can't see them doing more than this anytime soon.

Anyway, it's back to the code, and then the new hardware for the full-up speed tests.

Giving Up on “Clever” Constructor Usage – Go Simpler

Monday, August 9th, 2010

OK, I decided today to give up my clever solution to parsing these exchange messages into normalized messages. It was just too nasty. The basics seemed to work well enough, but if there were a value I wanted to set that wasn't in the exchange data feed - like which feed this was, then I had a real problem. If I used the "tag" list constructor, the value needed to be set, and the setters are protected because I want these guys to appear immutable to all the other parts of the system. So I had to make a friend relationship.

But because each of these codecs might generate any number of these normalized messages, I'd have a ton of friend statements in each class. That's just not supportable. At some point in the future, someone is going to forget to add one in, and then it's a mess as they try to figure out what's going wrong and fix it.

It's far easier to realize that each normalized message has a complete, typed, constructor, and all I need to do is to parse out the values from the exchange data stream, and then call that version of the constructor, and we're done. Nothing fancy. Not all tricky. Just plain and simple code.

While the other approach might have been nice to figure out how to skip values in the tag list, or put in functors for callbacks to the different date/time parsers for the different exchanges, it was all a bunch of mental gymnastics as it really didn't make the code any easier to read, any faster, or any more supportable. It was coder ego, and that's something I just hate to see.

So I simply cut out all the crud I'd made, and went back to a far simpler scheme, and the code looks much better for it.

Still Hammering Away at that Exchange Data

Thursday, August 5th, 2010

Today, in addition to a few meetings, I've been hard at work trying to work all the messages from one exchange data feed into my new ticker plant. The previous version of this project had over 300 messages - one per exchange message. I'm going for a far more minimalist design: if it's not a price or a trade, it's in a free-form variant message, and that's going to really help me in minimizing the number of messages people have to understand and deal with.

Yet there are always twists.

Today I realized that there are messages from the exchange that don't update their sequence number. OK, maybe they aren't critical to the function of the ticker plant, but I didn't want to throw them out at the UDP receiver. So I had to adjust the methods on my data source to allow for ignoring the sequence number checks. It made things a little more complex, but in the end, it'll be a better system.

That's where I'm at these days - taking what I think is a good idea and applying a real problem to it and seeing where it needs to stretch and fit. I'm hoping that I can get this all done in a few more days and then get to the real-world tests and verify that things are working as desired.

I've got my fingers crossed...

Struggling With Efficient Exchange Data Decoding

Wednesday, August 4th, 2010

Today I spent a lot of time trying to come up with a really nice way to parse the exchange data. It's not as simple as it seems. I should say that it's not simple to make something that doesn't take hundreds of lines of code comprised of structs, if/then/else and switch statements.

Typically, you're going to get a message with a fixed header. In that header, the full "type" of the message will be encoded. Then you lay another struct on the message and it reveals the other fields you can pick off. With each of these messages, you could be looking at from four to ten different "patterns" and therefore structs. This makes for a lot of code to really parse out the data.

Couple that with the switch statements to know what struct to apply, and the code gets very large, very fast. One upside to this scheme is that the execution of this code is very fast. So... I wanted something that was just as fast, but was far more compact in the 'lines of code' category.

What I decided to try was a "tagging" scheme where I don't attempt to make complete sense of all the data in the exchange record, but simply indicate where each value is, and how to decode it. For example, if I create the struct:

  typedef struct {
    char       type;
    uint16_t   pos;
    uint16_t   len;
  } variable_tag_t;

where the fields are type, position and length, then I can create an array of tags that indicates where the values I need are located in the data stream:

  static variable_tag_t[]  tradeTags = {
    { 'L', 0, 10 },
    { 'S', 10, 15 },
    { 'D', 25, 8 }
  };

and I can read this: I have a long int at position 0 that's 10 characters long, a string starting at position 10 for 15 characters, and a double at position 25 for 8. It's not bad, and I can see a lot of good with this idea.

I can create constructors on the messages that take one of these arrays and the data from the exchange and use it to extract the ivars from the data. This way, the order and count of tags is fixed by the message's needs, but the location and size of each value is dictated by the specifics of the exchange.

It seems like a decent idea, but I'm going to have to make several more messages and look at a few more exchange message definitions to make sure that it's really going to work out. I certainly like the compactness of the scheme. Some will argue that it's all hard-coded numbers, and that's not good - but how much different is this than a bunch of structs? Either is a hard-coded definition of the data organization in the exchange data. This just happens to be more direct.

We'll have to see how things work out.

Building Up my C++ Variant Class

Tuesday, August 3rd, 2010

Today I spent quite a bit of time really fleshing out my variant class today. I needed to have a lot more functionality in the code as it was going to be an integral part of the ticker plant I'm working on. I needed to write a bunch of tests, and each new test uncovered either a compiler issue - like needing a new version of a method for a different use case, or a real bug in the code, which I had to fix.

Overall, it was a pretty good day, but it was all spent making the class a lot more useful to the developer that would be using it. Which, of course, is me.