Saturday, June 28, 2014

Constant Urgency is a Symptom of a Hazardous Environment

When we use analogies like land-mines, fires, and various types of hell, we're referring to the hazards in our everyday work.  We may not be working with sharp objects or molten metal, but the anxiety, stress and exhaustion from working in a hazardous environment is very real.

From the outside, these problems are largely invisible.  So despite the hazardous work, we are usually expected to work without making mistakes.   We never have time to fix the hazards because there's always something more important to do.  Building the tools we need for safe development like failure recovery, diagnostic support, adequate logging and reliable deployment are often deferred in favor of more features.

Then when something explodes, as it inevitably will, high-risk heroics are required to save the day.  We work late nights and weekends repairing complex problems by hand and hope that nothing else goes wrong.

Instead of recognizing the symptoms of a serious problem, the long hours and heroics are often rewarded.  Fire-fighting, overtime, and last-minute hacks start to be expected.  Constant stress and exhaustion become the norm.

More people just add fuel to the fire.


Those that don't want to put in the long hours anymore are seen as not pulling their weight.  Frustration builds and the team gets burned out and the best developers start to leave.

The new guys just make things worse.  They don't know the software and the hazards to watch out for and they keep messing things up in the code.  We try to hold things together, but it's hard to get anything else done.  It becomes a full-time job just to keep the system from falling apart.

Management doesn't understand why productivity is so poor and tries to add more people to the work.  This just adds fuel to the fire.

Once this cycle gets started, it's hard to turn things around.  We get sucked into the problems, operating in a mode of constant urgency, and we don't want to see our project fail.  So we push ourselves to the limit of stress and exhaustion doing the best we can.  However, we're so busy reacting to all the things going wrong, there's no time to stop and fix the problems.  One more late night and a few hacks to get things working, but the cycle just doesn't end.

We knew better, but we did it anyway


The worst part about this is even when we know better, we do it anyway.

I remember one night in particular after working 60+ hour weeks for several months.  I checked in some code without running it at all and deployed my changes so I could test it in production.  I was so used to working under constant urgency, I had eventually thrown all my sense of principle out the window.

We had built out the delivery infrastructure and automated our release process from the beginning.  For a while, we were releasing every week; there were challenges, but for the most part things were going fairly well.  We had a major deadline coming up to support a new customer on our platform and investors had been promised it would happen by the end of the year.

The requirements meant drastically changing parts of the architecture and conquering some extremely difficult problems.  How long was it going to take?  We had no idea, but we did know we had better get to work!

We broke down the work and started chipping away at it, trying to do just enough unit testing to get by.  We paired on the more challenging parts and tried to parallelize the work to get it done as fast as we could.  We tried to integrate early, but there were so many problems.  The software produced weird results.  We just had to work through it.


We were caught up in the cycle


Some of us worked on testing and fixing, while others kept pushing along with the remaining features.  We knew we were headed down the path of a monstrous release, but we didn't seem to have any choice.  We worked an insane amount of hours troubleshooting problems just trying to get it stable.

The end of the year was rolling around and we finally got the software in production.  We thought the pain was finally over, but that was just the beginning.  We had no time to build out the infrastructure we needed to make changes safely, and our new users had a long list of complaints.  The pressure just never let up.

Every release it seemed like things would go wrong.  We'd work all weekend and be up late Sunday night trying to fix deployments that went wrong.  The data would be messed up.  Reports wouldn't be right.  We didn't really have a viable plan B.  The system was down, it took too long to restore from backup, we just had to fix it in production.


Something had to give...


We were so exhausted, but the urgency didn't end.  We were yelled at and threatened whenever things went wrong, but expected to continue the high-risk work.  How could they possibly give us bandwidth for work that wasn't part of the deliverables, when the project was already several months behind schedule?

We had poured so much of our time into the software and the people on the team were my friends.  We had great developers that had always been disciplined engineers and we all got sucked into the same trap.

Sometimes you just have to leave.  Working under threat and constant urgency makes great people do really stupid things.

Saturday, June 14, 2014

Designing Effective Teams

The same things that make for good software make for good team structure.  We need high cohesion within a team and low coupling between teams.  If people need high bandwidth communication across team structures to do their jobs effectively, the team structures are usually pretty dysfunctional.  Likewise if the members of a team don't have a need to talk to each other, they don't really operate as a team either.

Team structure is a design problem.  Developers can be quite good at it, once they start to look at it that way.  Designing the team structure around the architecture has a lot of benefits.  However, if you have a hairball interdependent architecture, you can't build a good team structure around it.  Trying to throw more people at the problem and artificially carve it apart is often where software organizations fail.  

Trying to go faster and throw more people at it often results in going *slower*. Teams get stuck in a trap of trying to police the code with reviews and there's no way to keep up.   The best resources can no longer be productive because they spend all their time reacting to the system that is busting at the seams.  Until the team learns a way to design the system in a way that *communication* can be scaled, leaders need to keep their foot off the accelerator pedal.   We need the time to invest in that critical learning.

I don't think it's a hands off, let the team figure it out kind of problem.  Organizational design is challenging problem and we need leadership to help figure it out.  But we need leaders that listen to their engineers, that know what to look for, and have an appreciation for the challenges of our craft.

Thursday, October 10, 2013

My new dev coaching...

Hey all,
Would you mind helping me to send this out to anyone you know that might be interested?  And a personalized endorsement of me would help too if you know me. :) 

I'm trying to get the word out... trying to build a developer coach-to-get-a-job program. :)  My new business experiment.

J

-------------
FREE Personal Development Coaching


Want to learn how to write cleaner code, or write more effective unit tests that aren't so annoying to work with?  How about learning a dev approach that will keep you feeling in control of your code, so the behavior stays predictable?

Sign up for personal one on one development coaching with me!  I'll meet with you, and teach you, and show you how.

The catch:

1) I'm only taking 5 people, because I need my sanity. :)

2) In exchange for being coached, you have to let me find you an awesome job that will help you grow further.  It's the awesomest place I know of to work in fact (other than New Iron of course).

3) Java takes preference, because I know it well.

4) And you have to be able to pass my interview. :)

Send me an email at janelle@newiron.com if you'd like to sign up. Since I have limited time, please say a few words about why this is important to you too, please.  And feel free to forward to anyone you think might be interested!

Wednesday, October 9, 2013

Lovin Life

I think my most productive days have been those roll out of bed and work days. Where you're just so into what you're doing it's the the first thing that pops into your head when you wake up. And you can't sleep because ideas keep popping into your head that you have to jot down. I just love the rush from being so excited about what tomorrow will bring, and ideas just coming together in the right way, its almost magical. I have an awesome life.

Two awesome breakthroughs: 

I figured out a working capacity model for shared business resources so that I could map capacity cost per client & revenue per client.  Talk about some amazing discovery opportunities.  What I've learned over the last week has been incredible.

And the other one, I figured out how to measure understandability and controllability in software so that the metrics matched my subjective evaluation.  I'm way excited about that - it's a huge breakthrough and the crux of my book and research effort!  I was so close, and could still make some things work ok, but everything seems to be falling into place.

I'm so incredibly excited.  

Wednesday, August 14, 2013

Got Database Pain?

I'm doing free brown-bag talks at companies around the community.   I have a lot of really great material after dealing with DB struggles in lots of different environments, and learning a lot along the way.  We'll go through patterns of common mistakes and how to reduce them, and strategies for making mistakes less costly when they do happen.   If you're interested, feel free to send me an email and we can schedule it!

Database CI: Practical Strategies for Reducing Database Pain

It's a common challenge - the database gets in the wrong state, the release is delayed, and the entire team is blocked and waiting for the centrally shared resource to be fixed.  Recovering from database mistakes can be quite painful.  Most of the tools available don't really solve this problem either - migrations focus on automating the deployment of the scripts, but don't help much in developing scripts that work to begin with.

But how can we reduce the number of mistakes? And how can we detect our mistakes as early and cheaply as possible?

In this presentation, we'll discuss strategies for reducing database costs with mistake-prone systems.  By adapting continuous integration principles to database work, we can drastically reduce the costs involved with database changes.  With practical examples and patterns for organizing SQL and database build automation, we'll cover strategies for many challenging issues:
  • Managing packages, procedures and views in the database
  • Changes (or mistakes) that can't be rolled back
  • High data volumes that make everything take longer
  • Systems that are hard to keep in your head
To register for this free brown-bag lunch presentation at your company, send us an email at contact@newiron.com and schedule a date! 

Saturday, August 10, 2013

Factors in DB Development Cost

I've been working on developing a new database CI tool suite, and was talking with a friend about his DB woes.  We talked about what problems he was having, where they occurred, and how long they were taking to resolve.   His database was really simple.  The scripts were pretty much always correct.  But some scripts might be forgotten, and his problems usually revolved around tracking down missing changes.

I realized my solution didn't fit his problems at all.  I had never worked in an environment that was all that simple.  There were always mistakes in the scripts.  And it was always painful to correct them.  When things went wrong, it was like the DB became the black hole of engineering hours, sucking away everyone's time.  We'd try to repair the mistake by hand.  Or if we couldn't figure out how to repair it, we'd have to restore from a snapshot of production.  And while all these repairs were going on, the engineering department was pretty much down.

Making mistakes was so painful.  And trying to use any of the database migration tools out there didn't seem to solve my problems.  In some cases, it even made them worse.  I always ended up resorting to rolling my own custom database tools.

So that got me thinking... how do you decide what kind of solution you need?

While there are many complex factors that drive DB development costs, there are two that seem to characterize the problem space quite well: frequency of mistakes, and cost of recovery.


When mistakes are rare, and recovery costs are cheap, existing migration tools are the perfect solution.  Most of the effort is spent in keeping databases up to date, and making sure all the scripts are deployed as the changes evolve.   Migrations are excellent at solving that problem, and handle it beautifully.

When mistakes are costly, migration tools can still be very helpful, but depending on how rarely mistakes occur, and how costly they are, migrations might not be enough.  Augmenting a simple migration tool might be a good option.

When mistakes are frequent, there's an entirely different kind of problem going on.  Developers aren't struggling with deploying the scripts, they're struggling to create correctly working scripts.  Migration tools tend to be intolerant of deploying and recovering from broken scripts.  And the tools can make recovery even more cumbersome by imposing additional constraints.

When frequent mistakes are also expensive to resolve... well, life is pain.   There's really only two strategies to a less painful life--figure out how to make fewer mistakes, or figure out how to make mistakes easier to recover from.  

I haven't found much help in this space in either the open source or commercial market, which is why I set out to try to fill that gap.  And I've built custom tools for solving similar problems several times over now, and had the opportunity to make a lot of mistakes.  Now I get to do something with all that learning, and hopefully help reduce some of the pain out there. :)


Friday, August 9, 2013

Wow how time flies...

Wow, has it ever been a while.  It was January, February, then suddenly it's August.  How the hell did that happen?

Some challenges came up at work that I had to jump in and get involved with, and most everything in my life started going on hold... including my book, and all the community stuff I wanted to do.  I got through writing chapter 4 back in February, and that's still where I'm at.   But I'm excited to announce having 2 uninterrupted days per week to focus just on going after making this happen.  I couldn't be more excited!

Since then the ideas started whizzing around in my head again.  After going to lunch with an old friend of mine, I was reminded of an idea I had been struggling with.  It had been at least a month, trying to make sense of something I didn't quite get, without words to describe what I thought was there.  But like a flash in my mind, he gave me a piece to the puzzle I was missing.  It was beautiful.

He was reading Thinking Fast and Slow, a book I read half of, and actually put down.  He said, "Association triggers from specific to general."  And that it had really struck him, "specific to general."

After reading On Intelligence, and working through the Art of Learning myself, I had been thinking about memory sequencing, and how it might impact our ability to recognize patterns.  Like when you see some messed up code, why is that sometimes an idea pops into your head about how to solve it, and other times it doesn't?  Why does this seem to happen more in some developer's heads than others?  Is this something that can be learned?  And can we learn how to do it faster?  This was the puzzle I was working on.

Specific to General.

We teach developers design patterns by handing them a book of design patterns.  A collection of "aha" moments from our predecessors.  But then armed with our new knowledge, we don't seem to have an ability to apply them.  When we see our own code, with our own problems, that flash of insight just doesn't happen.  Then for some, with experience, it happens.

Well, you could just say it's experience.  But couldn't we tailor the creation of the right experiences so that you would learn what you need to learn faster?

My friend had the key to my puzzle.  Specific to General.  The memory sequencing is critical, and the specific sequence of recognition is the opposite of what we teach.

If I want to have insight that leads to a design pattern, I need to experience specific problem instances that map to the more general pattern.  When I scan the code, the similarity of structural pattern to my specific memory should trigger recognition.

Going to have to pick up that book again I think. :)

Monday, December 3, 2012

My book is live!

I finally got my book effort officially kicked off now.  I'm planning on iterating through chapter development and getting the first early release done in January.  Also kicking off my new community group in January for Austin DIG and the Idea Flow project, a lot to do!

My awesome friend Wiley also helped me design the book cover. :)

http://leanpub.com/ideaflow


Saturday, July 21, 2012

Breaking Things That Work

I went to NFJS today, and went to a talk on "Complexity Theory and Software Development" by Tim Berglund.  It was a great presentation, but one idea in particular stuck out to me.  Tim described the concept of "thinking about a problem as a surface."  Imagine yourself surrounded by mountains, everywhere you look - you want to reach as high as you can.  But from where you are, you can only see potential peaks near by. Beyond the clouds and the mountain range obstructing your view, might be the highest peak of all.


With agile and continuous improvement come a concept of tweaking the system toward perfection.  But a tweak always takes me to somewhere nearby - take a step, if it wasn't up, step back to where I was.  But what if my problem is a surface, and the solution is some peak out there I can't see?  Or even if it is a peak I can see...  if I only ever take one small step at a time, I'll never discover that other mountain...

Maybe sometimes we need to leap.  Maybe sometimes we need to break the things that are working just fine.  Maybe we should do exactly what we're not "supposed to do", and see what happens.  

Now imagine a still pool of water... I drop in a stone, and watch the rings of ripples growing outward in response.  I can see the reaction of the system and gain insight into the interaction.  But what if I cast my stone into a raging river? I can certainly see changes in waves, but which waves are in response to my stone?  It seems like I'll likely guess wrong.  Or come to completely wrong conclusions about how the system works.

With all to variance in movement of the system - maybe it takes a big splash to in order to improve our understanding of how it works?  Step away from everything we know and make a leap for a far away peak?

Here's one experiment.  We've noticed that tests we write during development provide a different benefit when you write them vs when they fail and you need to fix them later.  How you interact with them and think about them totally changes.   So maybe the tests you write first and the ones you keep, shouldn't be the same tests?  We started deleting tests.  If we didn't think a test would keep us from making a mistake later, it was gone.   We worked at optimizing the tests we kept for catching mistakes. But this made me wonder about a pretty drastic step - what if you designed all code with TDD, then deleted all your tests? What if you added tests back that were only specifically optimized for the purpose of coming back to later?  If you had a clean slate, and you were positive your tests worked already, what would you slip in the time capsule to communicate to your future self?



Friday, June 8, 2012

What Makes a "Better" Design?

An observation... there are order of magnitude differences in developer productivity, and a big gap in between - like a place where people get stuck and those that make a huge leap. 

Of the people I've observed, it seems like there's a substantial difference in the idea of what makes a better design, as well as an ability to create more options. Those that don't make the leap tend to be bound to a set of operating rules and practices that heavily constrain their thinking. Think about how software practices are taught... I see the focus on behavior-focused "best practices" without thinking tools as something that has stunted the learning and development of our industry. 

Is it possible to learn a mental model such that we can evaluate "better" that doesn't rely on heuristics and best practice tricks? If we have such a model, does it allow us to see more options, connect more ideas? 

This has been my focus with mentoring - to see if I could teach this "model." More specifically a definition of "better" that means optimizing for cognitive flow. But since its not anything static, I've focused on tools of observation. By building awareness of how the design affects that flow, we can learn to optimize it for the humans. 

A "better" software design is one that allows ideas to flow out of the software, and into the software more easily.

Monday, June 4, 2012

Effects of Measuring

As long as measurements are used responsibly, not for performance reviews or the like, it doesn't affect anything, right?


It's not just the measurements being used irresponsibly - the act of measuring effects the system, our understanding, and our actions. Like a metaphor - metrics highlight certain aspects of the system, but likewise hide others. We are less likely to see and understand the influencers of the system that we don't measure... and in software the most important things, the stuff we need to understand better, we can't really put a number on. 

Rather than trying to come up with a measurement, I think we should try and come up with a mental model for understanding software productivity. Once we have an understanding of the system, maybe there is hope for a measurement. Until then, sustaining productivity is left to an invisible mystic art - with the effects of productivity problems being so latent, by the time we make the discovery, its usually way too late and expensive to do much about it. 

Productivity understanding, unlike productivity measuring, I believe is WAY more worth the investment. A good starting point is looking at idea flow.

Thursday, May 24, 2012

Humans as Part of the System

I think about every software process diagram that I've ever seen, and every one seems to focus on the work items and how they flow - through requirements, design, implementation, testing and deployment.  Whether short cycles or long, discreet handoffs or a collapsed 'do the work' stage, the work item is the center piece of the flow.

But then over time, something happens.  The work items take longer, defects become more common and the system deteriorates.   We have a nebulous term to bucket these deterioration effects - technical debt.  The design is 'ugly', and making it 'pretty' is sort of a mystic art.   And likewise keeping a software system on the rails is dependent on this mystic art - that seems quite unfortunate.  So why aren't the humans part of our process diagram - if we recognized the underlying system at work, could we learn how to better keep it in check?

What effect does this 'ugly' code really have on us?  How does it change the interactions with the human? What is really happening?

If we start focusing our attention on thinking processes instead of work item processes, how ideas flow instead of how work items flow... the real impact of these problems may actually be visible.  Ideas flow between humans.  Ideas flow from humans to software.  Ideas flow from software to humans.  What are these ideas?  What does this interaction look like?

Mapping this out even for one work item is enlightening.  It highlights our thinking process.  It highlights our cognitive missteps that lead us to make mistakes.  It highlights the effects of technical debt.  And it opens a whole new world of learning.

Thursday, April 12, 2012

Addressing the 90% Problem

If I were to try to measure the time that I spent thinking, analyzing and communicating versus actually typing in code, how much of the time would it be? If I were to guess, I'd say something at least 90%.  I wonder what other people would say? Especially without being biased by the other opinions in the room?

We spend so much of our time trying to understand... even putting a dent in improvement would mean -huge- gains in productivity. So...

How can we improve our efficiency at understanding?
How can we avoid misunderstanding, forgetting, or lack of understanding?
How can we improve our ability and efficiency at communicating understanding?
How might we reduce the amount of stuff that we need to understand?

These are the questions that I want to focus on... its where the answers and solutions will make all the difference.

Mistakes in a World of Gradients

I've been working on material for avoiding software mistakes, and have been searching for clarity on how to actually define "mistake."

I've been struggling with common definitions that are very black and white about what is wrong and considered "an error" or "incorrect".  Reality often seems more of a gradient than that, and likewise avoiding mistakes should maybe be more a matter of avoiding poor decisions in favor of better ones?

I like this definition better, because it accounts for the gradient in outcome, without missing the point.

"A human action that produces an incorrect or inadequate result. Note: The fault tolerance discipline distinguishes between the human action (a mistake), its manifestation (a hardware or software fault), the result of the fault (a failure), and the amount by which the result is incorrect (the error)."

 GE Russell & Associates Glossary - http://www.ge-russell.com/ref_swe_glossary_m.htm




Tuesday, April 10, 2012

What is a Mistake?

I've been trying to come up with a good definition for "mistake" in the context of developing software.

It's easy to see defects as caused by mistakes, but what about other kinds of poor choices? Choices that led to massive work inefficiencies?  And what if you did everything you were "supposed to" do, but still missed something, is that caused by a mistake?   What if the problem is caused by the system and no person is responsible, is that a mistake?

All of these, I think should be considered mistakes.  If we look at the system and the cause, we can work to prevent them.   The problem with the word "mistake", is it's quickly associated with blaming the who responsible for whatever went wrong.  Mistake triggers fear, avoidance, and guilt.  Which is the exact opposite of the kind of response that can lead somewhere positive.

Here's the best definition I found from dictionary.com:

"an error in action, calculation, opinion, or judgment caused by poor reasoning, carelessness, insufficient knowledge,etc. "

From this definition, even if you failed to do something that you didn't know you were supposed to do (having insufficient knowledge), its still a mistake.  Even if it was an action triggered by interacting parts of the system, but no one thing, its still a mistake.

But choices that cause inefficiencies? That seems to fall under the gradient of an "error of action or judgement".  If we could have made a better choice, was the choice we made an error? Hmm.

Sunday, April 8, 2012

A Humbling Experience

About 7 years ago, I was working on a custom SPC system project.  Our software ran in a semiconductor fab, and was basically responsible for reading in all the measurement data off of the tools and detecting processing errors.  Our users would write thousands of little mini programs that would gather data across the process, do some analysis, and then if they found a problem, could shutdown the tool responsible or stop the lot from further processing.

It was my first release on the project. We had just finished up a 3 month development cycle, and worked through all of our regression and performance tests.  Everything looked good to go, so we tied a bow on it and shipped it to production.

That night at about three in the morning, I got a phone call from my team lead. And I could hear a guy just screaming in the background.  Apparently, we had shut down every tool in the fab.  Our system ground to a screeching halt, and everyone was in a panic.  

Fortunately, we were able to rollback to the prior release and get things running again.  But we still had to figure out what happened.  We spent weeks verifying configuration, profiling performance, and testing with different data.  Then finally, we found a bad slow down that we didn't see before.  Relieved to find the issue, we fixed it quickly, and assured our customers that everything would be ok this time.

Fifteen minutes after installing the new release... the same thing happened.

At this point, our customers were just pissed at us.   They didn't trust us.   And what can you say to that? Oops?

We went back to our performance test, but couldn't reproduce the problem.  And after spending weeks trying to figure it out, and about 15 people on the team sitting pretty much idle, management decided to move ahead with the next release.  But we couldn't ship...

There's an overwhelming feeling that hits you when something like this happens.  A feeling that most of us will instinctively do anything to avoid.  The feeling of failure.

We cope with it and avoid the feeling with blame and anger.   I didn't go and point fingers or yell at anyone, but on the inside, I told myself that I wasn't the one that introduced the defect, that it was someone else that had messed it up for our team.

We did eventually figure it out and get it fixed, but by that time it was already time for the next release, so we just rolled in the patch.  We were extra careful and disciplined about our testing and performance testing, we didn't want the same thing to happen.

At first everything looked ok, but we had a different kind of problem.  It was a latent failure, that didn't manifest until the DBA ran a stats job on a table that crashed our system... again.   But this time, it was my code, my changes, and my fault.

There was nobody else I could blame but myself...  I felt completely crushed.

I remember sitting in a dark meeting room with my boss, trying to hold it in.  I didn't want to cry at work, but that only lasted so long.  I sat there sniffling, while he gave me some of the best advice of my life.

"I know it sucks... but it's what you do now that matters.  You can put it behind you, and try to let it go... or face the failure with courage, and learn everything that it has to teach you."

Our tests didn't catch our bugs.  Our code took forever to change.  When we'd try to fix something, sometimes we'd break five other things in the process.  Our customers were scared to install our software.  And nothing says failure more than bringing down production the last 3 times that we tried to ship!

That's where we started...

After 3 years we went from chaos, brittleness and fear to predictable, quality releases.  We did it.  The key to making it all happen, wasn't the process, the tools, or the automation.  It was about facing our failures.  Understanding our mistakes.  Understanding ourselves.  We spent those 3 years learning and working to prevent the causes of our mistakes.

Wednesday, March 28, 2012

What we REALLY Value is the Cost...

Today, someone in the community mentioned the idea of measuring "value points".  And the light went on... could this finally highlight our productivity problems?  It could be a totally dead-end idea, but its a hypothesis that needs testing.


When I thought "value points", I imagined a bunch of product folks sitting around playing planning poker judging the relative value of features. Using stable reference stories and choosing whether one story was more or less valuable than others. Might seem goofy, but its an interesting idea. My initial thought was that this would be way more stable over time than cost since cost varies dramatically over the lifetime of a project. And if it is truly stable, it might provide the missing link when trying to understand changes in cost over time.

For this to make sense, you gotta think about long term trends.  Suppose our team can deliver 20 points of cost per sprint. But our codebase gets more complex, bigger, uglier and more costly to change. Early on, we can do 10 stories at 2 points each.  But 2 years later, very similar features on the more complex code base require more effort to implement, so maybe these similar stories now take 5 points each and we can do 4 of them.  Our capacity is still 20 story points, but our ability to delivery value has REALLY decreased.

We often use story points as a proxy for value delivered per sprint, but think about that... We get "credit" for the -cost- of the story as opposed to the -value- of the story.   If our costs go up, we get MORE credit for the same work!

How can we ever hope to improve productivity if we measure our value in terms of our costs? How can we tell if a story that has a cost of 5 could have been a cost of 1? Looking at story points as value delivered makes the productivity problems INVISIBLE. It's no wonder that it's so hard to get buy in for technical debt...

What if we aimed for value point delivery? If you improved productivity, or your productivity tanked, would it actually be visible then? On that same project, with 20 cost points per sprint, suppose that equates to 10 value points early on, and 4 value points later.  Clearly something is different. Maybe we should talk about how to improve? Productivity, anyone? Innovation?

At least it would seem to encourage the right conversations...

Thursday, March 8, 2012

Does Agile process actually discourage collaboration and innovation?

Before everyone freaks out at that assertion, give me a sec to explain. :)

In the Dev SIG today, we were discussing our challenges with integrating UX into
development, and had an awesome discussion. I think Kerry will be posting some
notes. Most of the discussion though, went to ideas and challenges with
creating and understanding requirements, and the processes that we use to scale
destroying a lot of our effectiveness. The question we all left with, via Greg
Symons, was how do we scale our efforts while preserving this close connection
in understanding between the actual customer and those that aim to serve them?

In thinking about this, our recent discussions about backlog, and recalling past
projects, I realized some crucial skills that we seem to have largely lost. In
the days of waterfall, we were actually much more effective at it.

My first agile project was with XP, living in Oregon, and fortunate enough to
have Kent Beck provide a little guidance on our implementation. Sitting face to
face with me, on the other side of a half wall cube, was an actual customer of
our system, who had used it and things like it for more than 20 years. I could
sit and watch how they used it, ask them questions, find out exactly what they
were trying to accomplish and exchange ideas. From this experience I came away
with a great appreciation for the power of a direct collaborative exchange
between developers and real customers.

My next project, was waterfall. One of the guys on my team, wickedly smart, his
background was mainly RUP, and he just -loved- requirements process. What he
taught me were techniques for understanding, figuring out the core purpose,
figuring out the context of that purpose, and exploring alternatives to build a
deeper understanding of what a user really needs. Some of these were
documentation techniques, and others were just how you ask questions and
respond. I learned a ton. On our team, the customers would make a request, and
the developers were responsible for working with the customers to discover the
requirements.

With Scrum-esque Agile process, this understanding process is outsourced to the
product owner. As we try to scale, we use a product owner to act as a
communication proxy, and with it create a barrier of understanding between
developers and actual customers. Developers seldom really understand their
customers, and when given the opportunity to connect with them, the number of
discoveries of all the things we've been doing that could have been so much
better, are astounding.

I've done agile before on a new project, sitting in the same room with our real
users, understanding their problems, taking what they asked for and figuring out
what they needed, and also having control of the architecture, design, interface
and running the team to build it - the innovation of the project and what we
build was incredible. Industry cutting-edge stuff was just spilling out of
everything we did. And it all came out of sitting in a room together and
building deep understanding of both the goals, and the possibilities. This was
agile with no PO proxy. The developers managed the backlog, but really wrote
very little down... we did 1 week releases.

Developers seldom have much skill in requirements these days. And are often
handed a specification or a problem statement that is usually still quite far
from the root problem.

In building in this understanding disconnect, and losing these skills, are we
really just building walls that prevent collaboration and tearing down our
opportunities to innovate?

Manufacturing of a Complex Thought

Imagine that the software system is a physical thing. Its shape isn't really
concretely describable, its like a physical version of a complex thought. All
of the developers sit in a circle, poking and prodding at the physical thought -
adding new concepts, and changing existing ones.

Just like we have user interfaces for our application users, the code is the
developer's interface to this complex thought. Knowledge processes and tools
help us to manipulate and change the thought.

If I want to make a change, I need to first understand enough of the thought to
know how it would need to change. If I can easily control and manipulate parts
of the thought, and easily observe the consequences, its easier and faster to
build the understanding I need. Once I understand, I can start changing the
ideas, and again if I misunderstood something, it would be nice to know as early
as possible what exactly my mistake was, so that I can correct my thinking. If
there were no misunderstandings, the newly modified complex thought is then
complete.

In order to collaboratively work on evolving this complex thought, we must also
maintain a shared understanding - more brains involved increases the likelihood
of misunderstandings and likewise mistakes.

So with that model of development work, then think about all of the ideas and
thinking that have to flow through the process in order to support the creation
of this physical idea. Inventory in this context is a half-baked idea, either
sitting in the shelf, or currently in our minds being processed. These ideas
are what we manufacture, but since each idea has to be weaved into a single
complex thought - our tools that we use to control and observe the thought, the
clarity and organization of the thought, the size of the thought, all have a
massive impact on our productivity.

The tools are not the value, the ideas that get baked into this complex thought
are. All of the tooling is just a means to an end. We should strive to do just
enough to support the manufacturing of the idea.

If you think about creating tests from the mindset of supporting the humans and
these knowledge processes, a lot of what we do with both automated and manual
testing can clearly be seen as waste. An idea that is clear to everyone, for
example, is not one likely to cause misunderstandings and mistakes. We should
first aim to clarify the idea to prevent these misunderstandings. We should
then aim for controllable and observable, as these characteristics allow us to
come to understand more quickly. And when misunderstandings and mistakes are
still likely, we should then use alarms to alert us when we've made a mistake. 
False alarms quickly dilute the effectiveness of the useful alarms... thus
taking care in trying to point the human clearly to the mistake, and not raise
unnecessary false alarms is what makes effective tests.

Now think about things like code coverage metrics in this light. This metric
tends to encourage extreme amounts of waste. We forget all about the humans and
fill our systems with ever-ringing false alarms. We tend to only think about
tests breaking in our CI loop, but their real effect is their constancy in
breaking while we try to understand and modify this shared complex thought. 
With our test-infected mindsets, we quickly bury the changeability of our ideas
in rigidity, and lose the very agility that we are supposedly aiming for.

Friday, February 17, 2012

Fighting my way to agility - Part 5

Shrinking Release Size

Now that our batches were shrinkable, we could feasibly shrink our iteration and release sizes.  The ultimate test of whether you are -really- shippable is to actually ship.  If we could actually get our software in production, we could put the risk behind us for the changes we'd done so far.  If you don't actually ship, it was hard to -really- know if we were shippable.  We were also introducing more risk at a time to production - and troubleshooting production defects would usually take longer to diagnose and fix.

But our customers were just starting to trust us, and were still doing about a month of testing after our testing before they would feel safe enough to install something.   We asked if we could do releases more often, and the answer was pretty much... 'hell no.'  Rather than give up on trying, we figured out why there was so much pushback.  And solving those problems, whatever they might be, became our top priority.

We learned about all kinds of problems that we didn't even know we had.  Some were ticket requests that had been sitting in the backlog for years.  And others were just stuff that had to be done for every release that isn't really a big deal until we asked them to be done a whole lot more often.   There was a whole lot of pain downstream that we weren't even aware of.  And most of it was really just our problems rolling downhill - and completely within our power to fix.

Reducing the Transaction Cost of a Release

There's lots of different kinds of barriers to releasing more often.  Regardless of what yours are, getting the right people in the room, working together to understand the whole system goes a long ways. A lot of the stuff that seem like hard timeline constraints actually aren't.  Challenge your assumptions. 

So how could we relieve our customer's pains?

"Everytime we do a release, we lose some data" - we had no idea.  The system was designed to do a rolling restart, but there was a major problem.  During the roll over we had old component processes communicating with new ones.  In general the interfaces were stable, but there was still subtle coupling in the data semantics that caused errors.  Rather than trying to test for these conditions, we instead changed the failover logic so all of the data processing would be sticky to either the old version or new version and could never cross over.   This prevented us from having to even think about solving cross-talk scenarios.  We also created a new kind of test that mimicked a production upgrade while the system was highly live.  This turned out to be a great test for backward compatibility bugs as well.

"Its not safe to install to production without adequate testing, and we can't afford to do this testing more often" - Whether they found bugs or not while testing, was almost irrelevant.  Unless they knew what was tested and felt safe about it, they wouldn't budge.  They were doing different testing, with different frameworks, tools and people than we were, and unless it was done that way, it was a no go.  So we went to our customer site.  We learned about their fears and what was important to them.  We learned how they were testing and wanted testing to be done.   

We shared our scenario framework with them, code and all, and then worked with them to automate their manual tests in our framework.  We made sure their tests passed (and could prove it), before we gave them the release.  And likewise, we adopted some of their testing tools and techniques and at release time gave them a list of what we had covered.  We also started giving them a heads up about what areas of the system were more at risk based on what we had changed so they didn't feel so much need to test everything.  After we helped them to reduce their effort and just built a lot more trust and collaborative spirit with our customers, this was no longer an issue.

Editing the Scrum Rule Book - What process tools DID we actually use?

Since we weren't predictable and weren't using time boxes, we also threw out story point estimation, velocity and any estimation-based planning activities.  The theory goes that you improve your ability to estimate and therefore your predictability by practicing estimation.  I think this is largely a myth.

"Predictions don't create Predictability." This is one of my favorite quotes from the Poppendieck's book, Implementing Lean Software Development.   You create predictability by BEING predictable.  The more complex and far in the future your predictions, the more likely you are to be wrong - and way wrong.  So wrong that you are likely to make really bad decisions under the illusion that your predictions are accurate.   Its an illusion of control when control doesn't actually exist.  You can't be in control until you ARE controllable.   Predictability doesn't come from any process, its an attribute that exists (or doesn't) in the system.  Uncertainty is very uncomfortable.  But unless you face reality and focus on solving the root problem, nothing is ever really likely to change.

Burn downs we did use, but not until we were closer to wrapping up a release.  This was helpful in answering the 'are we done yet?' questions, the timing of which we used to synchronize other release activities.  There were tickets to submit, customer training to do, customer testing to coordinate etc.  We tried doing burn downs for the whole sprint, but since our attempted estimates were so wildly inaccurate - it wasn't helpful and more so harmful as input into any decisions.  The better decision input was that we really had no idea, but were trying to do as little as possible so done would come as soon as possible.  If a decision had to be made, we would try and provide as much insight as we could to improve the quality of the decision, without hiding the truth of uncertainty.  Although management never liked the answers, our customers were unbelievably supportive and thankful for the truth.