What Makes a Great Development Team

October 8th, 2007

I was talking to a vendor consultant that is here at the Shop helping us integrate their product into our systems. Happens all the time, we tell them what we have, they know what they need and we work together to hook up all the hoses and pipes so that their product fits as smoothly as possible into our systems, and in the end, becomes yet another one of our systems. So I was taking to this consultant letting him know what I did, as his product had to fit into my systems in a very smooth way.

We talked for about 2.5 hours... him asking me what my products were, how they worked, where the data flows were - nothing unusual. But then he asked me how many people were working on all this. I told him that there was a Team of seven on this part, and a few guys on this part, and there's a Team in the other part of the Bank for this part, but the rest of the whiteboard was me. He was shocked.

Which lead him into the question of what makes a really high-performance Team. Seeing as how I'm here to make this integration project succeed, I humored him and let him lead the conversation. His background, it seems, is in leading large-ish group efforts within the very large vendor's suite of applications. He said he was constantly trying to figure out what made Teams really special. Why did some groups seem to just walk in the light, and others were stuck constantly trying to stay off the Dilbert comics.

His take was experience and clarity of focus. He believed that if you had really experienced guys that knew the problem domain, who could look at an issue and easily distill it down to it's elemental key components, then that Team was going to be a winner. And the problem was having to deal with all the customer support and that meant that lesser talented developers had to deal with those issues. That lead to the problem of morale, and that lead to the fact that in any vendor group, 20% or less to the work and 80% are always the focus of the manager's efforts to make them productive.

I told him I thought he was all wrong. I also told him that if he had 80% of his developers that weren't contributing like the other 20%, he should fire them. And I went on to tell him why.

Maxim #1: People Rise to the Level of Your Expectations - as an educator both in one-on-one sessions and in small and large classes - both in the lab and the classroom settings, I believe this with all my heart. If you set your expectations high, people will achieve them. You have to do your part - you can't make them seem unattainable, you have to assist those willing to learn to get the knowledge, and you have to be fair, honest, and communicate these expectations, but if you do that, then people will rise to that level. I've seen it happen so many times, to me it's a universal truth.

So, for the manager that is trying to get the best out of the 80% - he needs only to expect the same from them as the others, and if those expectations aren't met, then those people have to leave the Team. Maybe they don't need to leave the company, but they need to know that they will not being doing development (cool work) with this group because it simply has expectations of each member that are beyond what this person is willing to do. In the end, you will loose a few, but the number won't be large, and even if it is, the Team will perform better because of it.

Maxium #2: Seek those with Commitment, and Everything is Possible - while many may disagree that this is true I'll put it another, less controversial way: If you find people that can acquire your commitment, they will be able to acquire anything else they need to see those tasks through to their ends. Education, experience, assistance - all these are available to the person with commitment. The consultant I was talking to spoke of experience and the ability to have a clear vision of the task at hand. Those are two qualities of a person with commitment, but those two qualities are not by themselves a complete indicator of commitment.

You can have people who are brilliant and do nothing. You can have people who see, and do not do. True, it's most often the case that people who can do these things are doing these things not because they are naturally good at them, but because they have worked to be good at them - natural talent or not. So while I understand what the consultant was saying, I think he was looking in the wrong place for it. Look for the traits that create and inspire the effects, as opposed to looking for the effects.

And this brings me to the core of the issue: Character. If you have it, you'll go far, and a Team with a critical mass of it will infect the others that might be a little low, and the result will be a Great Team. I've been on Teams where I'm the weakest person, and I worked very hard to not be the weakest person. I strove to be the person that could be relied upon to carry the day, if needed, but most times it was just a friendly competition to see who could do better - today. Tomorrow, the race starts again.

Gotta Watch Those Hotkey Combinations

October 6th, 2007

Yesterday I was playing with Skitch and FlySketch and saw that when I tried to do a screen snap in Skitch with the hotkey Cmd-Shift-5, I got a lot of problems and then didn't get the snap I wanted. The screen seemed to freeze, I couldn't get it to stop, and then I hit Enter, and I got a snapshot, but it wasn't what I wanted. I was convinced that this was just a conflict in the two applications - both of which use some kind of hotkey combination for their behavior. Not a terrible thing, but I don't like these conflicts as they sacrifice stability, so I picked Skitch over FlySketch, and wanted to give it a shake.

Well... overnight I thought more about it and wondered if it were possible to get the kind of things I liked in Skitch in FlySketch. So I flipped open the laptop and started looking at the preferences of FlySketch. When I saw that it's full-screen snapshot hotkey was Cmd-Shift-5 I knew that I'd found the problem. So since I have FlySktech set to pop-up on Cmd-Ctrl-F, I made the full-screen snapshot in FlySktech be Cmd-Ctrl-Shift-F - something easy to remember. Then the Cmd-Shift-5 would be Skitch's and there would not be a problem.

Or so I had hoped.

So I fired up Skitch again and along with the newly-configured FlySketch I tried the Skitch snapshot - Cmd-Shift-5. It worked!

Wonderful lesson to learn: Always check for hotkey conflicts.

There’s Nothing Quite Like Planning Ahead

October 5th, 2007

cubeLifeView.gif

OK, this weekend we have our disaster/recovery testing and while I've been ready for a while, several people are using today - yes, the last 24 hours before the test, to get things "set up" for the test. Now to be fair, there are some applications that don't need a lot of preparation, or there may be very limited expectations for some groups. But when those folks that decide to "check on things" at the last minute start to hammer me with requests for my systems to make their disaster/recovery testing go more smoothly, I don't really like it.

I have to agree that a lot of these are good things - mostly edge conditions related to what might happen in a disaster. And I have to say I'm happy to get these changes in no matter when, it's just that it would have been a lot nicer not to have to worry that I can get them fixed this afternoon, and had - oh... I don't know... maybe just another 24 hours... to get them in, tested and to production.

Alas, there are people that like to do things at the last minute. I, however, am not one of them.

Finally got a Skitch Invitation

October 5th, 2007

This morning I received my Skitch invitation. I read about Skitch from one of the new feeds I read daily, and the demo movie looked like it had a lot of nice features for image capture and sharing - though I'd be a lot more interested in the capture than the sharing, but there's stuff to share, and I guess they make their money off the adds they run on the viewing pages. Anyway, this morning I got he email saying I had received an invitation. So I decided to go back to the site and see if I was still interested.

Since I had put in the request many months ago, and I've started using FlySketch more, I wondered if I'd still think it was as "cool" as I did when I asked for the invite. Well... I have to say that it really is an amazing little tool. FlySketch is nice and small, and it can do a lot of things, but the ability to drag from Skitch to an app without saving (they save the image to /tmp and then 'drag' it in) and the ability to do more drawing than FlySketch is capable of makes Skitch a good upgrade from FlySketch.

Then when you look at the ability to put the images on their servers, allowing people to look at your sketches without you having to store them - well, that's just a nice bonus. Now I'm not sure I'd ever really trust a service that's free to keep my important images, but it's a place to throw stuff up there, and then decide if I need to keep a copy on my machine(s) later.

It's integration with OS X is really stunning and the mouse-over help is nice, but I hope as they mature the product it's able to be shut-off in the preferences. It works wonderfully well on my MacBook Pro - fast and snappy. I like the organization of the preferences even if it's a little against the Apple HIG. I have to say they did a nice job. I'll be running this and see where it'll take me.

UPDATE: it's a good idea to not have FlySketch and Skitch running at the same time. I've found some incompatibilities with the snapping in Skitch if you do. Since they both are about the same, I can see this as a limitation. I'm not really happy with it, but I can see it.

Debugging Replicated Database Problems

October 4th, 2007

database.jpg

Well... as I thought it might, the read-only copy of the instrument master database failed on me last night and while the primary is working fine, I feel it's necessary to be able to find a test case, or condition, where the replicated database fails so that I can give this to the team working on that project and they, in turn, can fix the underlying issue(s). I'm sure the local database admins will be be involved, as they have to be as we don't have that level of control over the servers and the machines. So, mauled by the sharks (from my previous post) I go back into the water trying to find the test case that will highlight the problem.

Last evening, the server was restarted at 5:49 pm, and the symbol set was divided into four groups of 889 underlyings and all four were sent out to the database proxy for loading. Typically, all four will finish within a few minutes of each other, but last night the first one finished at 17:55:12 and the second finished at 17:55:50 - but the third and fourth never finished. When I reconfigured the server to point to the primary, the four finished within 4 mins of each other - as they should. Clearly, there was something with the replicated database that was causing two of the loading threads to sit there waiting for data to come back. The question is, how to reproduce this?

It gets more of a quandary when you take into account that my development server started at 7:00 pm local time and it was fine using the read-only database - all four of it's loading threads finishing within a few minutes of each other. So there's something that's happening to the replicated copy between 5:50 and 7:00 pm that caused this problem, but it was gone by 7:00 pm.

I have a simple web page on the server's editor that allows me to look at the database operations that are being done in the code to see what the data is in the database and what's being retrieved. This has really helped a lot in the diagnosis of database issues like bad prices and missing key values. Yesterday, when we were having problems with the replication and the prices, I did have a few times when this page would not return all the data. Because it's a Perl script, it'd return what it had processed, but it would still act as if there were more to read (because there was), and yet nothing would come back. I'd love to be able to reproduce that for the guys.

Unfortunately, I haven't been able to. I have no tools at my disposal other than the requests I make. I'll keep hitting it throughout the day, but I don't hold out a lot of hope that this is going to point to anything conclusive. This leads me back to the same spot I was at yesterday - do I trust it? Today, however, the answer is different: No. I'll trust it when I have to trust it and not before. Since no one is really as concerned about this as I am, I'll stick to using the primary and see where the chips fall. There's no reason to risk production outages when all I've got to diagnose the problem is a few data loading scripts.

Depending on Other Groups

October 3rd, 2007

Interesting thing happened in the last 24 hrs. at work. We have a nice instrument master database - it holds all the instrument data as well as marks and such. Very nice. It's so nice that it's replicated to a read-only copy here in Chicago as well as a replicant in London for disaster/recovery purposes. It's great - except when the replication fails.

That's what happened last night. Normally, there's an explanation sent out about the "why" of why the replication failed. Changed tables, duplicate records - something caused the replication that normally works just fine to fail. But today I didn't hear what the cause of the break was. It appears to be fixed, but if we don't know why it failed, why can we assume that if it looks OK that it'll stay that way?

No reason to think so in my mind.

So when I ask the guys whose responsibility it is to maintain the database(s) "Is the read-only copy OK to use tomorrow?" and get a simple answer like "It should be good for tomorrow" I get a little nervous. What was the problem? Are we sure it's OK? I'm feeling like I'm on the beach and they're saying:

"Sure, Bob... go back in the water... pay no attention to the fins circling..."

"Oh, are we going in? No, not right now, just ate, don't want to get cramps... but you go right ahead... and splash a lot... here, hold this pork chop."

But since they are the ones responsible for the database(s), I'll listen to them and if I walk out of the surf a bloody, dismembered ghost of myself I'll say "Hey! You said it was OK!", and if there's fallout it'll be up to them, and not me, to answer those questions.

UPDATE: Well... as if I might have seen it coming, the read-only database was not ready for prime-time and I got the call at home after my server stalled on the restart due to waiting on the database. I switched it to the primary database and all was fine. Guess that means I'll have to figure out why.

New Acorn 1.0.2 Out Today

October 3rd, 2007

Today the guys at Flying Meat released Acorn 1.0.2 - there are a mostly bug fixes but there are a few new features too. Nothing that's super exciting right now, but it's nice to see the changes and even the little additions. I'm looking forward to 1.1 where Gus says there will be more significant additions like a 'Save for Web' where it'll try to figure a good way to save the image in a small size without compromising image quality.

And while we're on the subject of image editors, Photoshop Elements 6.0 is due out in 2008 for the Mac. They haven't really listed the features it'll have, but if it looks like PSE 6.0 for Windows it's going to be very different from what is PSE 4.0. It's been said that it looks like Lightroom, but having not seen that app, I can only say that the features look decent but the UI is very "enclosed" - not at all Mac-like. Very much like you'd run this app to the exclusion of other apps - no toolbars - one giant window. Very different. Not sure I like the change.

Interesting Bug in the Server

October 2nd, 2007

servers.jpg

This morning I got an email from one of the Hong Kong users about problems they were having with the theoretical values on OTC options in the server. While I didn't really get the proper picture from his email, the follow-up from a user in London provided the proper illumination to see what the problem was. Basically, when changing the volatility and/or dividends curve(s) for an instrument, the 'Open' greeks would be calculated properly and the 'Last' would not.

I dug into the code and saw that the 'Open' greeks are calculated no matter what, but the 'Last' are calculated only if they aren't in the cached data for that instrument already. This was the key - that the cached data was all wrong because it was generated based on the old curves, and with the new curves, all the cached data needed to be invalidated and recalculated.

Once I knew what I needed to do, it was just a matter of putting the code in-place to allow me to clear out the cache at the instrument level, then the underlying which would clear out all the derivatives' caches, etc. Then I had to put in the code to detect the change in the curves as read in from the editor, and putting it all together was pretty easy.

Once I had the code in-place, it was very easy to test and see that we are indeed getting the right greeks for changes in the curves. Also, I fixed a few issues in the editor so that it would not send the curves back to the server on an edit unless they had been edited by the user. This is going to help in simple efficiency as well as not making the server think things have changed when they haven't.

Added removeRows() to BKTable

October 1st, 2007

comboGraph.png

I had a new developer come up to me today and ask me if it were possible to remove a group of rows from a BKTable, and after a little bit of a sync on the terminology, it was clear that he wanted to delete a group of rows from the table based on some criteria. I was thinking about this and it seemed like a good idea even-though there's no way to do it now. So while he had to stick with iterating through the table and removing each one after testing it for it's applicability in the final results, I decided it was something I wanted to ask Jeff if he'd use it enough to add it to the class. Turns out, it's something that he'd like to see too. So I spent a little time today putting that into the BKTable.

I went with the idea that you'd provide this a JEP expression and it would either remove the rows that matched this expression or remove every row but those that matched this expression. Since I could use a lot of the same components that are in use in the filter table view this wasn't too hard. In fact, the effect is very similar to the filter table view, but in this case the removal is permanent and you can easily choose either set to remove.

After I got this in and tested I talked to Jeff and let him know that it'd be OK to have a list of a few things they wanted in BKit. He mentioned that his guys are suggested to come talk to me, with suggestions/problems so I guess there just aren't a lot of problems. OK.

Looking at Experimental Data

September 28th, 2007

Today I've spent a good deal of time working with a trader and a little app I wrote for their desk to pull some data from a service in a format that they can read into their applications. It's the kind of thing that I'll put a day or so into and they'll use it for a long time without modification simply because it's a data access component to them. No problem - in theory.

Yesterday, I worked on the socket communications for the clients of my market data server to change the way packets of data were read back from the server into the client's space. Rather than read a 2kB packet and then process it, I changed the code to read everything that was available at the socket, if anything was available. This meant that if we were receiving a 50kB message, we didn't read it in 25 chunks, we read it in one large chunk and then processed that. This made the processing of the data much faster because we didn't re-scan the first 2kB 25 times - we scanned it once. Now a 59kB packet is one thing, but some packets will be 1MB or more. Now we're talking significant savings in time.

Well... today I had to deal with someone that was not convinced that this was faster. In fact, they were convinced that it was slower. Given that they don't know the code, and only see that something has changed, I can understand their need for some kind of assurance that things have changed for the better. So what I did was to run two sets of trials: old versus new, five runs each, same data set to see what we'd get. Ideally, this will be a large enough sample set to be able to factor out the small (or large) variations in the access speed, network traffic, and other variables that you run into on large computer networked applications.

What I found out was that the time to gather the data from the source was somewhat variable, but the time it took to process the data once I had it was pretty controllable. The old way had times for this experiment from 43.2 to 43.7 sec - a pretty nice grouping, and the new way had times from 6.5 to 7.0 sec - again, a nice grouping. While the access times to get the data were much more variable, I made sure to include the access time as well as the processing time so that we could easily see where the time was spent.

Having spent more than a little time with experimental evidence myself, this was a nice sample size, and the breakout of the data made it clear where the variation was, and wasn't. Unfortunately, for this person, the data didn't say the same thing. They didn't see anything like that in the data. They saw the variation in the total time and said "See, the new is slower than the old here, so your changes hurt the system." When I tried to explain that if you looked at the breakout of the times, it was clear that the difference was the time spent in getting the data from the source (not in my control) and the processing time was nearly constant. But there was no convincing them.

For several hours I fought through this - more trials, asking what would convince them, etc. All the while, I'm thinking that these folks are extremely arrogant. It was only after I stepped back for a bit and looked at what they were saying that it hit me - they aren't arrogant, they're just horrible scientists.

Since their background is advanced degrees in science, I made the poor assumption that they actually were decent at reading experimental data and figuring out what the data is saying to them. After all, the job they have deals with experimental data every single day - it's called prices. Stochastic processes abound in this field, and it should have been second nature to these folks to read experimental data, but it's not. So me sense of frustration with their arrogance quickly turned into sadness about their lack of fundamental skills in this area of their work. Sad but true, I can't imagine trusting them with a dime of my money if they can't read experimental data like this.

As is so often the case, we carry in expectations to relationships that are sometimes far below the mark, and sometimes far above. It's all about really getting to know where the other folks are coming from and where their skills and weaknesses are. Now that I've properly calibrated myself to these folks, it'll be much easier to deal with them in the future.