Models of Reality
One of my general rules to constructing a system or even just modeling something for a system is “model reality.” Now, that might sound obvious but I never cease to me amazed with how many models and abstractions I work with are very much not even close to reality.
To model reality you have to first understand it.
The understanding part is, in my experience anyway, where the most trouble lies. It’s not a lack of desire or effort (well, not always). It’s rarely someone doing design deciding, “Yeah, let’s just make this thing I’m building very unlike what it represents.” The issue really presents itself when the designer or implementor believes they understand realities and then model their understanding instead of realities. The greater the gulf in harmony between what is and what is perceived, the greater the general pain for future you and your users.
In interviews I often ask the question, “If you had to chose between two similarly ranked candidates, where one was more technically skilled and really understands the tech stack and the other has a more modest technical background, but a deep understanding of the professional domain, who do you choose and why?” This is vague enough that it’s meant to be more a reasoning exercise than a question with a right answer. There really isn’t a right answer between the two hypothetical candidates. There are, of course, many wrong answers as to why one might choose one over the other.
In my own mind, I almost always favor the candidate with domain experience over technical wizardry. This stems from my own experience building systems in domains where I was extremely ignorant. I have stayed very much in the FinTech space for over a decade now and while I do not love the domain at all I have come to love what I can produce from having a deep understanding of its realities. That’s a factor of time and experience. There is no shortcut.
The thing about domains is that they are more a sum of their quirks rather than something well-documented and designed.
Moreover, the problems I have experienced with poor system design and long term problems related to those systems—outside of typical technical rot like lacking documentation and hurried implementation—are deep in the data modeling. These are the issues, especially in API design, that will haunt you like the specter of a tragic murder. They’re the reason companies consider the dreaded “v2” implementations of an API—which still never really separates you from the original problem.
Your mistakes proliferate more than you probably know. Bad domain modeling has to be worked around by integrators who actually know what they’re doing. Or worse, another company with engineers ignorant of their domain presume you knew what you were doing and they can rely on your models. Oops.
Technical issues, on the other hand, are typically not that hard to solve. I mean, it might take someone with some specialized technical skills to recognize and solve them but the path is often straightforward. Technical problems, unless they are abysmal, rarely become a problem until you’re starting to scale or deal with higher amounts of volume. This is almost always fine because volume—at least if your business is traditional enough that customers actually pay you—means cashflow and the ability to deal with with technical problems!
Another thing, these kinds of problems rarely require specialists and technical savants. At some point you have problems, improve your telemetry and instrumentation (or finally implement them into your system), identify bottlenecks, fix them, and have a good laugh. The true specialists are needed at crazy scale and if you’ve got crazy scale you’ve probably also got the crazy budget necessary to handle this problem.
Most of us—most of you—are unlikely to have the problems of crazy scale that you must also solve for. Even worrying about it is a huge waste of time and energy. Most of your problems of scale are going to be solved by learning how to set up proper database indexes and optimizing all I/O. These are just not that difficult to deal with.
You learn this when you see a company grow and it’s especially funny the first time. Tens of millions of rows in PostgreSQL ain’t nothin’ if you understand SQL and databases at even a journeyman level. To the uninitiated though, to someone who has traditionally worked on smaller systems, it sounds like a lot. These numbers cause engineers who don’t know and spend entirely too much time reading Medium drivel—written by engineers looking to pad their resumes and sound far more authoritative than they have any business being—to second guess simplicity in favor of nonsense.
That nonsense begets more nonsense and suddenly your entire organization, its systems, and its orchestration resemble a sort of corporate Rube Goldberg machine that is as intricate as it is unnecessary. Design patterns are exaggerated. Utterly preposterous paradigms become enshrined with religious reverence. Engineers engage in deep and long debates about literally everything in the system except how it serves the domain and the customer. Tech begins to exist for tech’s sake.
This is a mess. This is a mess that is usually built by people who are detached from the domain, from the customer, and from reality. Hyper-fixation on technical issues is a net negative.
So back to my question: the technical expert or the domain expert?
Personally, I favor domain knowledge over technical knowledge. Engineers rooted in their domain will almost always build superior products and be more in touch with their users. Even if the code itself is kind of ugly, the models and processes are more likely to be sensible and realistic. The ugliness can be refactored later by better programmers—or the same programmers 3-4 years later after they grew up a bit.
In a perfect world, a candidate or a lead designer is both. Poor fundamentals and the tech debt they ultimately create are still problems. It’s just that, in my experience, when the original model is more sound and the data and processes are consistent with how non-technical people in the domain operate, the system will be inherently more flexible. Problems that are mostly technical are much easier to Fix Later™.
That’s really the focus of reality and how we stay connected with it—the users. In corporate and engineering circles, I want people to think of me like Tron. I want them to say things like:
That’s Tron. He fights for the users.
You can’t fight for the users if you don’t understand them. You can’t fight for the users when you’re disconnected from them. You can’t fight for the users when you have no idea what they actually want out of the software you’re building.
You have to understand the domain. You have to understand your own product. If your work and understanding do not stretch beyond the checkboxes cooked up by a manager in your project management system, your output will be poor—which is the case with software in general.
Don’t get me wrong, a lot of software is bad because the companies designing it view users are vessels to be squeezed for maximum profit. This generates the same problem though. The makers are ultimately disconnected from the users. User need is filtered through sales and product. Engineers, who often do not interact customers at all, are building in vacuums and echo chambers. Worse, left to their own, engineers will often fixate on technical issues—not system issues which are, surprisingly, related far less than one would presume!
I used to force all the engineers on my team to spend time in the support queue. This is a great way to keep engineers from using their support staff as a convenient buffer to reword, “You’re doing it wrong.” If your UI makes it too easy to “do it wrong” then maybe the problem isn’t the user.
Wouldn’t you rather work on a product that your users enjoy and gladly pay for rather than something that feels more like Stockholm Syndrome? Do you honestly want to work on a system that is begrudgingly tolerated?
Domain knowledge and dare I say empathy for the user goes hand in hand with making well-designed systems. Neither of these facets is particularly technical either. Your technical skills aren’t unimportant, but if they’re not grounded in the former ideals, the end result is what one would expect—the product of a disconnected engineer making disconnected decisions that are rooted in tickets rather than realities.
Let me tell you, that’s no way to live.
The Consequences of Reliance
My career in FinTech has taken some strange, almost incestuous twists. The end result of which has allowed me to watch a few once startups become much bigger operating companies. I’ve seen first hand how far my own ignorance has spread. It’s interesting. It’s horrifying. I almost wish I didn’t know how widespread my own mistakes could manifest.
In early 2014 I was hired to build a crowdfunding portal. It was the definition of “a job.” I didn’t care about the space at all. I didn’t understand the space at all. No problem though because for the first time in five years I’d be but a guy on a team that would just tend to the technical matters. After years at other startups as the engineer, it seemed like it would be a nice break. In fact, my specific Ruby experience—meaning I wasn’t just a “Rails guy” and actually had a basic understanding of databases outside of an ORM—is what landed me the job.
Six months in, I was the suddenly the de facto CTO and our funding platform had pivoted into a FinTech providing APIs for compliance and money movement. This was probably the single biggest period of technical and professional growth in my career. It’s amazing what you can do when you have no choice and have to keep the lights on.
Our engineering “team” was reduced to two, and later we got a third who would form the true core of the company for three years. I was not only lead engineer and system designer but also handled IT in the office, infrastructure, upper tier support, documentation, and was often the main voice with our most important customers. It was amazing in so many ways. The lack of separation between me and the business was as educational as it gets.
And not to brag, but I kicked ass. I grew. I learned so much. I also got to see how early ignorance plagued future me and customers over the years.
I left in 2020 after watching the CEO completely forget about who actually built the company once investor cash started pouring in1. I got to see the engineering team grow vastly and even pay a ton of outsourced engineers. It’s surreal to watch a budget ballon, names multiply, processes pop up, and yet see very little substantive growth in output. All it seemed to do was increase our ability to build roads to nowhere and completely lose focus on our own business. (For the record, this is where my hatred of product departments really took hold.)
A couple years later I would be hired by a new startup that was adjacent. I was hired specifically because of my experience and connections in the space. This new company was funded by and serving a customer who had been one of the biggest customers of my prior venture. I knew some of their staff. I knew their product. As a new partner with a different role, I’d learn all kinds of new things and have a more intimate view of their system.
To my shame and horror, as I had to help export data into our new system, I realized tons of their modeling errors were familiar. When they were a newborn startup, they had relied on my prior company’s APIs for their entire business. Remember that part where I said:
I didn’t understand the space at all.
The early API was designed and put together very much in the “didn’t understand” era. The CEO—who was the primary product designer—took the approach of, “You’re a smart guy. You’ll figure it out.” And I did, but it took a lot of time (and hilariously, a lot of secret conversations with industry people we worked with). By the time I had anything to call sea legs, we were all living with fundamental weaknesses in the data model. Since our API was our main product, we couldn’t “just fix” many of our most deep-rooted problems2.
Unintended Poisoning of the Well
The early data models used by our API were underdeveloped and full of things that were just wrong unless you took the Obi-wan “certain point of view” point of view. They were inflexible and short-sighted. This all stemmed from me not understanding the space and our CEO’s belief that it didn’t matter. I was just the engineer. I didn’t need to “get it.” I just needed to follow his own “specs”3 to the letter.
Regardless, our APIs—warts and all—were the backbone of many businesses in the space and were absolutely necessary for them to carry out their business. Despite their warts, this system worked and I cared a lot about it. Working with our actual customers and the engineers doing the integrations gave me a very personal perspective on it all. After all, how would I feel integrating an API that was sometimes… nonsense? (I feel rage, by the way. That’s the emotion.)
This customer I was working with at my new stop had been one such business in its earlier life and six or so years later was also suffering from modeling problems it inherited from coupling tightly and ignorantly to my modeling problems! They had copied some of our most broken and incomplete concepts to the letter. Why? Their engineers knew the space even less than I had and this company was all about the 7 layer Product burrito between engineers and vendors. No one there was having secret meetings with the guy running the broker-dealer to try and make sense of the madness. “Oh, the API vendor does it X way. That must be How It Is Done™.”
My new gig was similar in terms of business model, but the foundational design was arranged by current me with near a decade of experience. As turns out, you learn a few things over eight to nine years. So here I am, helping a company map terrible, broken, misunderstood data from their system that was based almost entirely on my own former self’s terrible, broken, and misunderstood data model to one that reflected actual reality—reality that I learned the hard way I might add!
It was surreal to see this full circle. If you have ever integrated with a vendor, you probably presume they know what they’re doing. It’s probably much more like the above. There’s probably a lesson here about education, tight coupling, and other things. I leave it to you to divine that part for yourself.
What I wish to impart on you, however, is that your mistakes and ignorance probably have farther reaching consequences than you believe.
Technical Problems Are Usually the Easy Ones
The problems with our API that were difficult to work around were the modeling issues. The technical problems like scale, algorithmic efficiency, and things of that nature? Pieces of cake by comparison.
I was so worried about size and scale when I started. “Web scale” as a concept was really getting popularized around then. Twitter was the poster child for it. It made me scared of size without realizing that the gulf between “large and fine with a single database for writes” and needing to really scale is absolutely gigantic. After about two years as we grew, had plenty of traffic and paying customers, we were handling it fine and I wasn’t employing anything magical. Ruby on Rails and a single database was barely breaking a sweat most days.
We would run into issues though. One particular customer decided to poll our most inefficient endpoint like every minute—including paginating all the records—and crippled our database and services. After first getting them to stop we decided we needed to build a better notification system than “a lot of customers are pissed off” and get some better instrumentation in general.
This took some time, but again, not that much and the work was pedestrian anyway. These were settled science that were in no way industry specific. That’s why technical problems tend to be easier to deal with. If anything, the hardest part is the vast amount of choice you have in solving them.
Moreover, the actual endpoint? The problem was a collection of n+1 queries that resulted in hundreds of queries when called for each request. This error had been there for years at this point. The actual fix took me about four hours to track down, patch, write tests for, and deploy once we realized there was a problem. I added a few indexes to the database and changed the query in the code to use some joins rather than rely on the ORM to lazy load relationships. This is not even remotely interesting work.
The customer could actually resume polling4 at their prior interval after that and did so for a couple more years despite constantly promising webhook integration was “right around the corner.”
I saw this sort of thing a lot. When I was more novice, these things scared me. I believe inexperienced engineers tend to worry far more about technical than domain and design issues. When faced with them myself though, they were relatively easy to deal with. We handled them with a skeleton crew5 of three guys for close to as many years.
Let me emphasize, we’d been operating with this inefficiency for years without a major issue. We fixed it in hours.
Users, Domain, and Design
The only way to really know a domain in the business sense is to interact with your users and customers or do so indirectly through those in your organization who do. This is why I have traditionally enjoyed my time with support and sales—when I wasn’t wearing those hats myself! I have participated in may joint calls and meetings with customers. Teams should encourage this rather than wall off engineering.
I think product departments have grown very much around this idea that programmers are somehow unable to communicate with normal people. This is entirely counter to everything I have ever experienced with any company. The number of savant-like engineers I’ve encountered is honestly zero. I know they exist in the wild, but their myth is powerful in ways that have mostly negative effects.
Outsourcing also creates huge issues here. When you outsource with developers who are neither linguistically nor culturally familiar with your customers and live in a time zone that’s several hours different, you’re gonna have problems. Frankly, I’ve never understood outsourcing at all. It’s a money pit.
Users and customers aren’t always synonymous. Your internal staff is very much a user of your system in some way. Smart companies understand that heavy internal quality and automation are essential and a borderline super power. Companies run by the more prolific Homo Erectus exclusive C-suite invest as little time and resources as possible on internal matters and view them as “cost centers.” (Yes my friends, capitalism and MBAs are the gifts that keep on giving.)
For those who work primarily on APIs, your users are other engineers. You have absolutely no excuse in this realm for not understanding your users, their needs, their expectations, and what would simply make their lives easier. You know exactly what they want because you want it. You’ve probably integrated multiple APIs that have taken weeks longer than they should have because each new system is a chance to experience someone else’s irrationality that is so insane you couldn’t even predict it. APIs function as their own sort of Cthulhu Mythos.
This is all a very roundabout way of saying:
- Data modeling is the heart of system building.
- Modeling reality should be your North Star.
- Understanding realities is a combination of understanding the domain and its users.
- Empathy for those users will result in a superior system, even where it has technical deficiencies.
- The leadership in your organization likely understands none of this. (This is a running theme.)
In closing, remember this: Tron is awesome6. Be like Tron. Fight for the users. Model their realities, quirks and all. Everything you build will be better for it.
-
The danger of money is leaders often think they can now go purchase experts and discard the people who actually built the system and made it successful enough to land that cash in the first place. I’ve seen this cycle play out a few times at a few companies. Also, a funny thing about “experts” in a field is they regularly have no idea about an industry beyond their last job. This is a huge problem if the scale of the current company does not match their last stop. ↩
-
This is where a lot of people go to the dreaded v2. And we did. The new API was much better and… guess how many clients adopted it that were already firmly entrenched in v1? The point where you try to force v2 on customers is the point where competitors become an option because a new integration with you is probably not much more difficult than looking elsewhere. If your relationship is already strained, guess how that ends? When you don’t really care about your customers, you’re mostly banking on their own inertia. ↩
-
These things were hilarious. I would walk into the office and find a half page print out from Microsoft World with some red pen mark up. There was one where he wrote “etc.” in terms of things that needed to be handled. I had to explain, “You do realize I’m the details guy, right? My whole job lives in etc. You’re going to need to be a bit more specific.” ↩
-
This complete aside belongs down here. I do not understand polling. I just don’t. We had webhooks. I only poll when I have literally no other option but in my experience engineers start with polling because they think webhooks or other asynchronous means are “harder.” I have never found this to be true, especially with the related downsides of polling. This is one of the most baffling “shortcuts” that I see taken over and over again. Just… why!? ↩
-
I’m sure some readers are already like, “How did you not have…?” And the answer is, three guys running everything including sales calls means you do only what you must. And while I extoll the virtues of good software development, YAGNI is also worth understanding. The question is always, “Even if we do need it at some point, is it a big deal now?” This was probably the best point made by The Lean Startup. While I don’t recommend the book at all actually because it could have about 1/10th the size, I did enjoy their many examples of “we gotta” followed by, “oh wait, we don’t gotta.” ↩
-
It’s worth noting that Tron helped bring down a thieving criminal senior executive vice president—perhaps the most corporate title imaginable. I’ll just leave that right there. ↩