<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Mohit Ranka</title><link href="https://www.mohitranka.com/" rel="alternate"/><link href="https://www.mohitranka.com/atom.xml" rel="self"/><id>https://www.mohitranka.com/</id><updated>2026-07-16T10:00:00+05:30</updated><entry><title>Near-real-time is a product promise, not a Kafka cluster</title><link href="https://www.mohitranka.com/blog/near-real-time-is-a-product-promise/" rel="alternate"/><published>2026-07-16T10:00:00+05:30</published><updated>2026-07-16T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2026-07-16:/blog/near-real-time-is-a-product-promise/</id><summary type="html">&lt;p&gt;“Near-real-time” is often used as a synonym for “we bought a streaming stack.” On GTM data platforms, that confusion is expensive. The product promise is about &lt;strong&gt;whether a decision-maker can trust a number in time&lt;/strong&gt;—not whether an event log is busy. At LinkedIn, the Enterprise Data Platform (EDP) was …&lt;/p&gt;</summary><content type="html">&lt;p&gt;“Near-real-time” is often used as a synonym for “we bought a streaming stack.” On GTM data platforms, that confusion is expensive. The product promise is about &lt;strong&gt;whether a decision-maker can trust a number in time&lt;/strong&gt;—not whether an event log is busy. At LinkedIn, the Enterprise Data Platform (EDP) was intended as the governed center for GTM datasets. BI teams on Power BI and Tableau still ran on Hadoop-era batch paths. They already had data. What they did not have was a reason to treat EDP as the path that made GTM analytics true, timely, and durable as legacy pipelines were marked for deprecation. That gap is a product-promise gap.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/near-real-time-is-a-product-promise.jpg" alt="Abstract illustration of a clock and data streams feeding a product dashboard" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Near-real-time is a promise about the age of truth a consumer can act on—not a badge for a streaming cluster.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;What consumers actually bought&lt;/h2&gt;
&lt;p&gt;When sales, marketing, or BI stakeholders ask for better data timing, they rarely mean “please introduce a new consumer group.” They mean things like: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The dashboard does not contradict the operational truth for long.&lt;/li&gt;
&lt;li&gt;A GTM metric used in a weekly motion is not secretly a stale extract.&lt;/li&gt;
&lt;li&gt;There is one place to stand when two numbers disagree.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Those are promises to &lt;strong&gt;named consumers&lt;/strong&gt;. For us, the critical early consumers included BI engineering and the GTM partners who lived in their dashboards. If those consumers do not move, the platform’s internal latency graphs are cosplay. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;A definition that forces honesty&lt;/h2&gt;
&lt;p&gt;I use a boring definition: &amp;gt; For dataset &lt;em&gt;D&lt;/em&gt; and consumer &lt;em&gt;C&lt;/em&gt;, the promise is that the &lt;strong&gt;age of usable data&lt;/strong&gt; at the moment &lt;em&gt;C&lt;/em&gt; acts stays inside an agreed bound—under normal conditions—with a clear story when it does not. Unpack it in platform language: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Named dataset&lt;/strong&gt; — a sales or GTM entity teams recognize, not “the lake.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Named consumer&lt;/strong&gt; — Power BI workbook owners, Tableau extracts, revenue analytics—not “downstream.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Usable&lt;/strong&gt; — passes the governance and correctness bar, not merely “row arrived.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agreed bound&lt;/strong&gt; — may start as a milestone (“on EDP before deprecation date, with acceptable query latency”) before it becomes a polished SLO.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure story&lt;/strong&gt; — what the business does when the path is wrong: freeze a pipeline, pin a version, staff a war room—not “we’ll check the dashboard Monday.”&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you cannot name &lt;em&gt;D&lt;/em&gt; and &lt;em&gt;C&lt;/em&gt;, you are not ready to sell near-real-time. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Why “we can already get the data” kills platform promises&lt;/h2&gt;
&lt;p&gt;The BI objection was rational: existing Hadoop-based jobs still produced outputs. From their seat, migration was risk without immediate upside. So the competing product was not another vendor. It was &lt;strong&gt;the legacy batch path that still worked&lt;/strong&gt;. Near-real-time platforms lose to “good enough yesterday” until:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;the old path has an end date,&lt;/li&gt;
&lt;li&gt;the new path is staffed for adopters,&lt;/li&gt;
&lt;li&gt;query performance after cutover is somebody’s on-call problem.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;We treated those as part of the promise design. EDP engineers helped migrate. Connectors reduced friction into Power BI and Tableau. Deprecation deadlines made dual-running finite. Performance work made “governed” not mean “slower.” Streaming could have been part of some paths. It was not the adoption strategy.&lt;/p&gt;
&lt;h2&gt;Hidden decisions that are actually the product&lt;/h2&gt;
&lt;h3&gt;Where is the system of record?&lt;/h3&gt;
&lt;p&gt;If EDP and a legacy pipeline disagree, which number is allowed to win in a QBR deck? Until that is explicit, faster pipelines just produce faster arguments. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h3&gt;Is the promise contractual or best-effort?&lt;/h3&gt;
&lt;p&gt;Deprecating Hadoop paths is a contractual move: the company is choosing a continuity posture. Best-effort “please try EDP” will lose to local convenience forever. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h3&gt;Who pays for fan-out?&lt;/h3&gt;
&lt;p&gt;Every BI team and every GTM dataset multiplies support surface. Platforms that treat every new consumer as free eventually stop being able to keep any promise. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h3&gt;What happens during catch-up and cutover?&lt;/h3&gt;
&lt;p&gt;Migrations create windows where two truths coexist. The product promise must cover the dual-run period—or you will invent tribal knowledge in Slack. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;The shape that worked: phases, not a big-bang bus&lt;/h2&gt;
&lt;p&gt;The useful architecture picture was organizational as much as technical: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Executive buy-in&lt;/strong&gt; — GTM data consolidation as strategy, not a side project.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MVP on high-value sales datasets&lt;/strong&gt; — prove the promise where pain is visible.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Expand across sales and revenue consumers&lt;/strong&gt; — widen the interface only after the path works.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Broader GTM standardization&lt;/strong&gt; — reduce snowflake pipelines.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance and optimization&lt;/strong&gt; — make the default path the boring path.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That sequence is how you keep a freshness/correctness promise while the org is still learning to trust the platform. “Put everything on a stream in quarter one” is usually a way to buy complexity before you have consumers. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Questions I ask before approving “let’s go real-time”&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Which decision improves if this dataset gets younger—and who makes that decision?&lt;/li&gt;
&lt;li&gt;What is the current path’s delay, really (including BI extracts and cache TTLs)?&lt;/li&gt;
&lt;li&gt;What dies when we succeed—the legacy job, or only our spare time?&lt;/li&gt;
&lt;li&gt;Who is staffed to migrate the first consumers?&lt;/li&gt;
&lt;li&gt;What query latency is acceptable in the tools people actually use?&lt;/li&gt;
&lt;li&gt;What do we measure weekly that a VP would recognize?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If the answers are all technology choices, the promise is not ready. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Near-real-time is not a badge for an architecture review. It is a &lt;strong&gt;promise about time, truth, and behavior when truth is late&lt;/strong&gt;. EDP’s lesson was blunt: the company does not get that promise when a platform exists. It gets that promise when critical consumers—here, BI on GTM data—run on the governed path, and the old batch defaults are allowed to end. Build streams when they earn their keep. Build the consumer promise first.&lt;/p&gt;</content><category term="Blog"/><category term="data-platforms"/><category term="distributed-systems"/><category term="reliability"/></entry><entry><title>The platform team’s real job is interfaces</title><link href="https://www.mohitranka.com/blog/platform-teams-real-job-is-interfaces/" rel="alternate"/><published>2025-11-13T10:00:00+05:30</published><updated>2025-11-13T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2025-11-13:/blog/platform-teams-real-job-is-interfaces/</id><summary type="html">&lt;p&gt;At LinkedIn, the Enterprise Data Platform (EDP) was meant to be the centralized way GTM teams managed and consumed datasets. On paper, that is a clear platform charter. In practice, a platform is only real when its &lt;strong&gt;interfaces get adopted&lt;/strong&gt;—including by teams that already have a path that “works …&lt;/p&gt;</summary><content type="html">&lt;p&gt;At LinkedIn, the Enterprise Data Platform (EDP) was meant to be the centralized way GTM teams managed and consumed datasets. On paper, that is a clear platform charter. In practice, a platform is only real when its &lt;strong&gt;interfaces get adopted&lt;/strong&gt;—including by teams that already have a path that “works.” The hard case was BI. Power BI and Tableau teams were still living on Hadoop-based pipelines scheduled for deprecation. They could already get data. EDP was strategically important and still optional in their week. Without them, EDP could not become the source of truth for GTM analytics no matter how good the internals looked. That is an interface problem, not a cluster problem.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/platform-teams-real-job-is-interfaces.jpg" alt="Abstract illustration of connecting building blocks and interface ports between teams" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Platforms earn leverage when the interfaces other teams stand on are clear, adoptable, and owned.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;The interface was “how BI gets governed data”&lt;/h2&gt;
&lt;p&gt;When platform teams say interface, they often mean API shape or event schema. Here the consumer-facing interface was broader: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How does a BI engineer get a trusted dataset into a dashboard workflow?&lt;/li&gt;
&lt;li&gt;Who pays the migration cost?&lt;/li&gt;
&lt;li&gt;What happens to the old path, and when?&lt;/li&gt;
&lt;li&gt;Is query performance acceptable the week after cutover?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;EDP’s storage format and pipeline elegance did not answer those questions. Until they were answered, BI had no reason to reorder priorities. &lt;strong&gt;Lesson:&lt;/strong&gt; product teams do not consume your architecture diagrams. They consume time-to-success, predictability, and whether the platform team shows up when the path is rocky.&lt;/p&gt;
&lt;h2&gt;Adoption failed as an org problem first&lt;/h2&gt;
&lt;p&gt;This was easy to misread as stubbornness. It was incentives. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;BI teams were measured on analytics delivery, not on platform migration.&lt;/li&gt;
&lt;li&gt;Legacy pipelines still produced outputs.&lt;/li&gt;
&lt;li&gt;Migration looked like unfunded work with downside risk (latency, rework, surprise breakage).&lt;/li&gt;
&lt;li&gt;EDP’s long-term governance story was real—and still abstract compared to this quarter’s dashboards.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Engineering alignment inside the data org was necessary and insufficient. Without a business framing, “please adopt EDP” is a favor request. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Executive sponsorship changed the type of conversation&lt;/h2&gt;
&lt;p&gt;I partnered with senior leaders on the data and BI side so EDP adoption was not a side quest. The useful reframe was not “modernize for us.” It was: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Hadoop paths were going away.&lt;/li&gt;
&lt;li&gt;Continuing to depend on them was a &lt;strong&gt;continuity risk&lt;/strong&gt;, not a neutral default.&lt;/li&gt;
&lt;li&gt;EDP was the consolidation path for governed GTM analytics—not a nice-to-have alternate store.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That shift moved the discussion from technical preference to operating risk. Platforms that cannot get that sentence said out loud usually stall in permanent pilot mode. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;We lowered the price of yes&lt;/h2&gt;
&lt;p&gt;Even with sponsorship, asking BI to self-fund a migration would have failed slowly. The decision that mattered operationally: &lt;strong&gt;EDP engineers would take migration load&lt;/strong&gt;—connectors, pairing, performance work—not only publish docs and wish for pull requests. Concretely, that meant: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Pre-built paths into Power BI and Tableau so “get data from EDP” was not a research project.&lt;/li&gt;
&lt;li&gt;Shared work on query performance so cutover did not mean slower dashboards.&lt;/li&gt;
&lt;li&gt;A migration toolkit and recurring office hours to kill blockers in public, early.&lt;/li&gt;
&lt;li&gt;Alignment with sales and marketing stakeholders who depended on the outputs, not only the producers.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is platform-as-product without the theater: the job-to-be-done was “keep GTM analytics working on a governed foundation,” and we priced the platform to make that job rational. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Deprecation is part of the interface&lt;/h2&gt;
&lt;p&gt;Enabling a new path without disabling the old one is how companies collect platforms. The adoption plan included the unglamorous half: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Clear deprecation milestones for legacy Hadoop pipelines.&lt;/li&gt;
&lt;li&gt;Time-bound dual running where needed.&lt;/li&gt;
&lt;li&gt;A definition of done that included &lt;strong&gt;turning things off&lt;/strong&gt;, not only turning EDP on.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Governance is not a slide about ownership. Governance is whether the abandoned path still quietly feeds production dashboards six months later. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;What changed&lt;/h2&gt;
&lt;p&gt;Within a concentrated push—on the order of a quarter for the BI migration motion—Power BI and Tableau usage moved onto EDP for the scoped GTM paths we targeted. Redundant pipeline surface area could be decommissioned. Freshness and governance improved because fewer competing “sources of truth” were allowed to linger. I care less about the trophy phrasing than about the mechanism: &lt;strong&gt;sponsorship + funded migration + forced deprecation + BI-shaped interfaces.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Where EMs earn their keep on platform teams&lt;/h2&gt;
&lt;p&gt;The technical work was real. The EM job showed up in places that never appear in a system design doc: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Spending team capacity on someone else’s migration so the company’s data model could converge.&lt;/li&gt;
&lt;li&gt;Holding a deprecation line when temporary extensions would have been easier.&lt;/li&gt;
&lt;li&gt;Making sure performance issues after cutover were our problem, not a gotcha that punished adopters.&lt;/li&gt;
&lt;li&gt;Translating platform risk into language executives will prioritize.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Special cases still existed—BI tools always have them—but they were absorbed into connectors and support rituals, not infinite private forks. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Platform teams love infrastructure. Infrastructure is not the job. The job is to define, evolve, and defend &lt;strong&gt;interfaces other teams can build on&lt;/strong&gt;—including the incentives, migration labor, and deprecation schedule that make those interfaces real. EDP did not become the GTM source of truth when it launched. It became the source of truth when BI could succeed on it, and the old paths were allowed to die.&lt;/p&gt;</content><category term="Blog"/><category term="engineering-leadership"/><category term="platforms"/><category term="developer-tooling"/></entry><entry><title>When I still choose a relational database</title><link href="https://www.mohitranka.com/blog/when-i-still-choose-a-relational-database/" rel="alternate"/><published>2025-03-13T10:00:00+05:30</published><updated>2025-03-13T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2025-03-13:/blog/when-i-still-choose-a-relational-database/</id><summary type="html">&lt;p&gt;In 2013 I wrote &lt;a href="https://www.mohitranka.com/blog/rdbms-vs-nosql/"&gt;RDBMS vs. NOSQL?&lt;/a&gt; as a pushback against fashion. The fashion changed costumes—document stores, wide-column, NewSQL, "Postgres is fine," "everything in the lakehouse"—but the underlying mistake did not: &lt;strong&gt;picking a datastore from a blog post instead of from access patterns and failure modes.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/when-i-still-choose-a-relational-database.jpg" alt="Illustration of an ordered data table foundation beside scattered documents" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;A relational …&lt;/span&gt;&lt;/figcaption&gt;&lt;/figure&gt;</summary><content type="html">&lt;p&gt;In 2013 I wrote &lt;a href="https://www.mohitranka.com/blog/rdbms-vs-nosql/"&gt;RDBMS vs. NOSQL?&lt;/a&gt; as a pushback against fashion. The fashion changed costumes—document stores, wide-column, NewSQL, "Postgres is fine," "everything in the lakehouse"—but the underlying mistake did not: &lt;strong&gt;picking a datastore from a blog post instead of from access patterns and failure modes.&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/when-i-still-choose-a-relational-database.jpg" alt="Illustration of an ordered data table foundation beside scattered documents" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;A relational spine is often the boring default until access patterns and failure modes force a different shape.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;p&gt;This is the sequel I would write to myself: when I still choose a relational database in 2026, and when I do not. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;The default remains boring on purpose&lt;/h2&gt;
&lt;p&gt;My default for a new product backend is still a managed relational database (usually PostgreSQL) with: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Clear schema ownership&lt;/li&gt;
&lt;li&gt;Migrations as code&lt;/li&gt;
&lt;li&gt;Backups and point-in-time recovery you have actually restored&lt;/li&gt;
&lt;li&gt;Connection pooling and boring observability&lt;/li&gt;
&lt;li&gt;A plan for read scale that is not "wishful replicas"&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Why? Because most products are &lt;strong&gt;transactional workflows with relationships&lt;/strong&gt;: users, accounts, permissions, orders, configurations, audit trails. Relational databases are extraordinarily good at that shape. They also match how humans reason about correctness. If your primary problems are multi-row invariants, ad hoc query flexibility, and operational maturity, starting elsewhere is often self-inflicted difficulty.&lt;/p&gt;
&lt;h2&gt;When relational is the right call&lt;/h2&gt;
&lt;p&gt;I lean relational when several of these are true: &lt;strong&gt;1. Strong invariants matter more than infinite write scale.&lt;/strong&gt;&lt;br&gt;
Money, entitlements, identity bindings, "exactly one active X per Y." If the business invariant is relational, fighting the model is expensive. &lt;strong&gt;2. Query patterns are still evolving.&lt;/strong&gt;&lt;br&gt;
Early products change questions weekly. A well-modeled schema with indexes beats a write-optimized store that makes new questions painful. &lt;strong&gt;3. The team’s operational muscle is SQL-shaped.&lt;/strong&gt;&lt;br&gt;
A perfect paper architecture with zero operators is worse than a known system with runbooks. Skills are part of architecture. &lt;strong&gt;4. Multi-entity transactions simplify the product.&lt;/strong&gt;&lt;br&gt;
Saga forests can be correct. They are also a tax. If a single-node or lightly clustered RDBMS can hold the transactional core, keep the core small and sharp. &lt;strong&gt;5. You need ecosystem gravity.&lt;/strong&gt;&lt;br&gt;
ORMs, migration tools, BI access, hiring, incident folklore—relational ecosystems are deep. That is not marketing; it is time-to-recovery.&lt;/p&gt;
&lt;h2&gt;When relational becomes the wrong center&lt;/h2&gt;
&lt;p&gt;I move work &lt;em&gt;out&lt;/em&gt; of the primary OLTP database when: &lt;strong&gt;1. Write throughput or state size exceeds honest vertical + read-replica plans.&lt;/strong&gt;&lt;br&gt;
Not vanity metrics—measured saturation, vacuum pain, replica lag that product feels, backup windows that scare you. &lt;strong&gt;2. Access patterns are append-heavy and query-narrow.&lt;/strong&gt;&lt;br&gt;
Event logs, high-volume telemetry, massive multi-tenant time series. Different stores exist for reasons. &lt;strong&gt;3. Fan-out reads need specialized shapes.&lt;/strong&gt;&lt;br&gt;
Search, graph traversal at scale, feature stores, geospatial at high QPS—often better as &lt;strong&gt;derived systems&lt;/strong&gt; fed from a transactional core. &lt;strong&gt;4. Availability topology requirements outgrow your relational operator model.&lt;/strong&gt;&lt;br&gt;
Some multi-region active-active stories are possible with modern relational systems; many are still research projects wearing production clothes. Be honest about the topology you can run. &lt;strong&gt;5. Schema flexibility is real, not aesthetic.&lt;/strong&gt;&lt;br&gt;
Truly heterogeneous documents with little shared query structure can be a poor fit. "We might need flexibility later" is not a requirement.&lt;/p&gt;
&lt;h2&gt;The architecture that aged best for me&lt;/h2&gt;
&lt;p&gt;The pattern I trust: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Transactional system of record&lt;/strong&gt; in relational (or something with equal invariant strength).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explicit integration events&lt;/strong&gt; when other systems must react—versioned, owned, documented.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Derived read models&lt;/strong&gt; for specialized query paths.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Analytical systems&lt;/strong&gt; for heavy aggregation—do not abuse OLTP as a warehouse.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clear rules for dual writes&lt;/strong&gt; — preferably avoid them; if not, make correctness visible.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The relational database is not the universe. It is often the &lt;strong&gt;spine&lt;/strong&gt;. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;A related lesson from platform service boundaries&lt;/h2&gt;
&lt;p&gt;Not every “monolith vs microservices” fight is a datastore fight—but the same judgment applies. On an EDP self-serve portal, we had a month-long freeze over whether new UI workflows had to live entirely in a monolithic backend or split early for independent iteration. The useful answer was hybrid: keep &lt;strong&gt;core platform capabilities&lt;/strong&gt; where integration and invariants dominate; split &lt;strong&gt;fast-changing workflows&lt;/strong&gt; where deploy independence pays for the seam cost; put &lt;strong&gt;metadata&lt;/strong&gt; in a system designed for lifecycle, not in accidental dual writes. That is the same muscle as choosing a relational spine: &lt;strong&gt;optimize for ownership and correctness at the center, derive or split at the edges, refuse fashion-driven rewrites that stop delivery.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What changed since 2013&lt;/h2&gt;
&lt;p&gt;A few updates to my older self: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Managed Postgres (and peers) are much better.&lt;/strong&gt; Failover, backups, and scaling knobs improved. That raises the bar for "we need NoSQL to be reliable."&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;JSON columns done carefully&lt;/strong&gt; cover many former document-store arguments—without abandoning transactions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CDC became mainstream.&lt;/strong&gt; Turning relational truth into streams is a product pattern, not a science fair.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Distributed SQL exists&lt;/strong&gt; and can be right—but it is still a specialist choice with specialist costs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lakehouses ate a chunk of analytics&lt;/strong&gt;, which is good: stop pretending your OLTP DB is a lake.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What did &lt;em&gt;not&lt;/em&gt; change: &lt;strong&gt;NoSQL is not a personality.&lt;/strong&gt; It is a set of tradeoffs around consistency, query power, operational complexity, and scale. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Decision checklist I actually use&lt;/h2&gt;
&lt;p&gt;Before approving a non-relational primary store: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Which queries and invariants are first-class in v1?&lt;/li&gt;
&lt;li&gt;What is the 12-month data size and QPS with ugly margins?&lt;/li&gt;
&lt;li&gt;What does multi-row correctness look like, and who enforces it?&lt;/li&gt;
&lt;li&gt;How do we migrate schema or access patterns six months in?&lt;/li&gt;
&lt;li&gt;Who is on-call, and what have they operated before?&lt;/li&gt;
&lt;li&gt;Can we explain the choice in one paragraph without vendor slogans?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If step 6 requires a conference talk, the choice may still be right—but it needs more proof. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;I still choose a relational database when I want &lt;strong&gt;correctness, evolvable queries, and operational boredom&lt;/strong&gt; around the core of a product. I choose other systems when the access pattern or scale curve makes that boredom impossible. The winning move is rarely "pick a side." It is &lt;strong&gt;keep a sharp system of record, derive aggressively, and refuse to let fashion rename your requirements.&lt;/strong&gt; If you want the older, more argumentative version of this stance, it is still here: &lt;a href="https://www.mohitranka.com/blog/rdbms-vs-nosql/"&gt;RDBMS vs. NOSQL?&lt;/a&gt;. The industry moved. The need for judgment did not.&lt;/p&gt;</content><category term="Blog"/><category term="data-platforms"/><category term="databases"/><category term="architecture"/></entry><entry><title>Freshness SLOs: the metric product teams actually feel</title><link href="https://www.mohitranka.com/blog/freshness-slos-the-metric-product-teams-feel/" rel="alternate"/><published>2024-07-11T10:00:00+05:30</published><updated>2024-07-11T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2024-07-11:/blog/freshness-slos-the-metric-product-teams-feel/</id><summary type="html">&lt;p&gt;API latency SLOs changed how we run services. GTM data work taught me the sibling idea the hard way: product teams do not experience your job-success chart. They experience a dashboard that is late, wrong, or disagrees with another “official” number. At LinkedIn, BI teams on Power BI and Tableau …&lt;/p&gt;</summary><content type="html">&lt;p&gt;API latency SLOs changed how we run services. GTM data work taught me the sibling idea the hard way: product teams do not experience your job-success chart. They experience a dashboard that is late, wrong, or disagrees with another “official” number. At LinkedIn, BI teams on Power BI and Tableau could already pull data from Hadoop-era batch paths while the Enterprise Data Platform (EDP) was supposed to become the governed center for GTM datasets. Pipelines can be “green” while the business still feels stale or fragmented truth. That feeling is a freshness and trust problem—even when nobody is paging on consumer lag.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/freshness-slos-the-metric-product-teams-feel.jpg" alt="Illustration of a freshness gauge on an operations dashboard" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Product teams feel freshness as trust in the number—not as job-success charts on a platform dashboard.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;Latency is not freshness&lt;/h2&gt;
&lt;p&gt;A fast dashboard query can still serve an extract that is hours old. A successful batch job can still leave sales and marketing acting on last night’s world. Latency asks: how long did this request take?&lt;br&gt;
Freshness asks: &lt;strong&gt;how old is the truth this decision used?&lt;/strong&gt; If you only measure the serving API, you will celebrate the wrong layer.&lt;/p&gt;
&lt;h2&gt;Define freshness in the consumer’s language&lt;/h2&gt;
&lt;p&gt;For EDP-shaped work, a useful definition was: &amp;gt; For a named GTM dataset and a named BI consumer, how old is the data at the moment someone uses it in a dashboard decision—and is it the governed path? That forces specifics: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dataset&lt;/strong&gt; — a sales or GTM entity people recognize.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consumer&lt;/strong&gt; — a Power BI or Tableau workflow, not “analytics.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Usable&lt;/strong&gt; — on the platform you claim is source of truth, not a leftover pipeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Age&lt;/strong&gt; — including batch boundaries and BI-side refresh behavior, not only platform ingest time.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You can formalize that into an SLO later. First you need a sentence a BI lead and a platform EM would both underline. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Promises that change behavior (not charts for their own sake)&lt;/h2&gt;
&lt;p&gt;On the BI migration, the promises that changed behavior were not abstract percentiles on a poster. They looked like: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Continuity:&lt;/strong&gt; legacy Hadoop paths had deprecation dates—staying put was a risk posture, not a neutral default.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Funded cutover:&lt;/strong&gt; EDP engineers helped migrate; adopters were not asked to donate a quarter of calendar alone.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Acceptable query performance after switch&lt;/strong&gt; — so “governed” did not mean “slower dashboards.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Finite dual-running&lt;/strong&gt; — two truths were a migration window, not a lifestyle.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Call those SLOs if your org has the discipline. Call them &lt;strong&gt;operating promises&lt;/strong&gt; if you are still early. Either way, something has to hurt when the promise breaks—attention, prioritization, or the ability to keep the old path alive. I will not pretend every dataset had a polished error budget. What we had was executive visibility, migration milestones, and a definition of done that included turning legacy paths off.&lt;/p&gt;
&lt;h2&gt;What to measure along a real path&lt;/h2&gt;
&lt;p&gt;For a BI consumer, time and trust hide in stages: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Source systems and upstream delay  &lt;/li&gt;
&lt;li&gt;Platform ingest and validation on EDP  &lt;/li&gt;
&lt;li&gt;Dataset readiness / governance checks  &lt;/li&gt;
&lt;li&gt;BI tool refresh or extract behavior  &lt;/li&gt;
&lt;li&gt;Cache and workbook-level assumptions&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Teams often discover the villain is not the fanciest processor. It is an extract schedule, a dual pipeline, or a cutover that never finished. Instrument the path your consumers actually use. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Correctness sits beside freshness&lt;/h2&gt;
&lt;p&gt;Fresh wrong data is worse than slightly stale right data. During migration, dual sources create a special failure mode: two numbers, both “recent,” different owners. Correctness signals that mattered in practice: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Is this dataset on the governed path?  &lt;/li&gt;
&lt;li&gt;Are legacy feeds still quietly serving production workbooks?  &lt;/li&gt;
&lt;li&gt;Do critical fields null out or drift after cutover?  &lt;/li&gt;
&lt;li&gt;Can someone name the system of record when a QBR fights itself?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Freshness without governance optimizes for speed of confusion. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;How we introduced the promise for BI&lt;/h2&gt;
&lt;p&gt;The sequence that worked was not “roll out SLO framework company-wide”: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Pick the consumer class already on the critical path (BI for GTM).  &lt;/li&gt;
&lt;li&gt;Make adoption a leadership priority with a continuity narrative.  &lt;/li&gt;
&lt;li&gt;Staff the migration so the new path is cheaper than it looks.  &lt;/li&gt;
&lt;li&gt;Put dates on deprecation.  &lt;/li&gt;
&lt;li&gt;Fix performance issues as platform bugs, not user error.  &lt;/li&gt;
&lt;li&gt;Only then talk about tightening time bounds dataset by dataset.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If you start with twenty datasets and a perfect taxonomy, you will get a wiki. If you start with one embarrassing consumer journey, you might get a habit. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;The objection we actually heard&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;“We can already get the data.”&lt;/strong&gt;&lt;br&gt;
Yes. That is why platform-only arguments fail. Answer with end-of-life for the old path, labor for the new one, and proof that cutover does not degrade the dashboard. &lt;strong&gt;“Batch is fine.”&lt;/strong&gt;&lt;br&gt;
Sometimes it is. Then write a batch promise (“available by time T for the weekly motion”) and still own correctness and deprecation. Batch without a promise is how shadow pipelines live forever.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Product teams feel freshness as trust in the number. Platforms earn that trust when named consumers run on a governed path, old paths can die, and someone is accountable when the age of truth is wrong. Whether you brand it SLO or operating promise matters less than whether the company can point to &lt;strong&gt;dataset + consumer + bound + owner&lt;/strong&gt;—and whether missing it changes next week’s work.&lt;/p&gt;</content><category term="Blog"/><category term="data-platforms"/><category term="reliability"/><category term="observability"/></entry><entry><title>Saying no as a platform EM without becoming the villain</title><link href="https://www.mohitranka.com/blog/saying-no-as-a-platform-em/" rel="alternate"/><published>2023-11-08T10:00:00+05:30</published><updated>2023-11-08T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2023-11-08:/blog/saying-no-as-a-platform-em/</id><summary type="html">&lt;p&gt;Platform engineering managers do not run out of good ideas. They run out of capacity to say yes to every reasonable request without wrecking the shared system. The skill is not blunt refusal. It is &lt;strong&gt;no with a path&lt;/strong&gt;—specific enough that partners can execute, firm enough that your team …&lt;/p&gt;</summary><content type="html">&lt;p&gt;Platform engineering managers do not run out of good ideas. They run out of capacity to say yes to every reasonable request without wrecking the shared system. The skill is not blunt refusal. It is &lt;strong&gt;no with a path&lt;/strong&gt;—specific enough that partners can execute, firm enough that your team is not a free consulting desk. Two nos from EDP work at LinkedIn taught me more than any generic prioritization framework.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/saying-no-as-a-platform-em.jpg" alt="Illustration of a fork in the road with one path closed and an alternate route open" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;A useful platform “no” names the constraint and funds a path—not a silent block.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;Why platform “no” feels personal&lt;/h2&gt;
&lt;p&gt;Product and BI partners are graded on outputs this quarter. Platform teams are graded on leverage, reliability, and whether the company still has one data model next year. When you decline an unfunded migration or a rewrite-shaped preference, it can sound like indifference. Sometimes that critique is fair. Often it is a missing alternative. A villain blocks silently. A partner names the constraint and funds a way through.&lt;/p&gt;
&lt;h2&gt;Principles I actually use&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Company throughput beats local speed.&lt;/strong&gt; A one-week special case that creates a permanent support branch is not kindness.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A delayed honest yes beats a fake yes.&lt;/strong&gt; A Jira key without staffing is a lie.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tradeoffs go in writing the same day.&lt;/strong&gt; Memory is political; notes are kinder.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Everything below is those three principles in concrete form. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Case A — No to “just adopt the platform”&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Request (implied):&lt;/strong&gt; BI teams on Power BI and Tableau should move to EDP because it is the strategic GTM data platform. &lt;strong&gt;Reality:&lt;/strong&gt; They could already get data from Hadoop-based pipelines. Migration looked like unfunded risk. EDP could not become the source of truth without them. &lt;strong&gt;The no:&lt;/strong&gt; No to a pure mandate without labor—“adopt EDP” as a favor to the platform team. &lt;strong&gt;The path:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Executive sponsorship that framed legacy pipeline deprecation as &lt;strong&gt;continuity risk&lt;/strong&gt;, not taste.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;EDP engineers assigned to migration work&lt;/strong&gt;, not only documentation.&lt;/li&gt;
&lt;li&gt;Connectors into Power BI and Tableau.&lt;/li&gt;
&lt;li&gt;Query performance work so cutover did not punish adopters.&lt;/li&gt;
&lt;li&gt;Tooling, office hours, and &lt;strong&gt;hard deprecation milestones&lt;/strong&gt; so dual-running ended.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That package is a no to magical adoption and a yes to an expensive but real interface change. Within a concentrated push (about a quarter for the BI motion we scoped), the migration stuck and legacy surface area could shrink. &lt;strong&gt;Pattern:&lt;/strong&gt; If the consumer has no incentive, your “no” is to unfunded asks; your “yes” is to change incentives and price.&lt;/p&gt;
&lt;h2&gt;Case B — No to both pure extremes in an architecture fight&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Request (implied):&lt;/strong&gt; Pick monolith &lt;em&gt;or&lt;/em&gt; microservices for an EDP self-serve portal—each side sure the other choice was malpractice. &lt;strong&gt;Reality:&lt;/strong&gt; The disagreement went public, ownership collapsed, and delivery froze for roughly a month. &lt;strong&gt;The no:&lt;/strong&gt; No to a binary holy war. No to indefinite debate. No to “loudest critique wins.” &lt;strong&gt;The path:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Structured design review with &lt;strong&gt;written criteria&lt;/strong&gt; (scalability, maintainability, speed, ownership).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Time-boxed POCs&lt;/strong&gt; from both approaches instead of slide wars.&lt;/li&gt;
&lt;li&gt;A neutral senior engineer in the room.&lt;/li&gt;
&lt;li&gt;Explicit coaching on ownership and influence in 1:1s—not only technical arbitration.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;hybrid decision&lt;/strong&gt;: core platform capabilities stayed integrated with the EDP backend; more dynamic portal workflows could be separate services; metadata lifecycle centralized (we used DataHub) rather than re-invented.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Execution resumed on the order of a week after the decision landed; the portal followed on a months-long path with real adoption. &lt;strong&gt;Pattern:&lt;/strong&gt; Sometimes the EM’s no is to false dichotomies. The funded path is a hybrid with proofs. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Make yes expensive in the right way&lt;/h2&gt;
&lt;p&gt;Temporary exceptions will exist—dual pipelines during migration, transitional architecture branches. Price them: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Time-bounded&lt;/strong&gt; (deprecation date, not vibes)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Owned&lt;/strong&gt; (named team for breakage)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Visible&lt;/strong&gt; (on a list leadership can see)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Removable&lt;/strong&gt; (exit criteria written down)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Free, quiet exceptions are how platforms drown. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Roadmaps are how you say no at scale&lt;/h2&gt;
&lt;p&gt;One-off negotiation does not survive GTM scope. The EDP sales/GTM program needed an explicit sequence: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Buy-in and prioritization  &lt;/li&gt;
&lt;li&gt;MVP on high-value sales datasets  &lt;/li&gt;
&lt;li&gt;Expand across sales/revenue consumers  &lt;/li&gt;
&lt;li&gt;Broader GTM standardization  &lt;/li&gt;
&lt;li&gt;Governance and optimization&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That roadmap is a machine for “not yet.” Without it, every dataset is an emergency and every emergency is a yes. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Protect the team without hiding behind them&lt;/h2&gt;
&lt;p&gt;“The team is busy” is weak if you cannot show the math. Better: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Here is committed platform work (migration staffing, deprecation, portal seams).&lt;/li&gt;
&lt;li&gt;Here is what we will not staff this quarter.&lt;/li&gt;
&lt;li&gt;Here is the escalation if the business wants to reorder.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Take heat in partner forums so individual engineers are not negotiating company priority alone. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Saying no as a platform EM is stewardship of shared constraints. On EDP, the nos that mattered were: &lt;strong&gt;no unfunded adoption&lt;/strong&gt;, and &lt;strong&gt;no architecture theater that freezes delivery&lt;/strong&gt;. The yeses were expensive on purpose—engineers on migration, connectors, deprecation, POCs, hybrid seams. If partners can see the path, you are not the villain. You are how the company keeps one platform instead of twelve.&lt;/p&gt;</content><category term="Blog"/><category term="engineering-leadership"/><category term="platforms"/><category term="management"/></entry><entry><title>Identity systems fail socially before they fail cryptographically</title><link href="https://www.mohitranka.com/blog/identity-systems-fail-socially/" rel="alternate"/><published>2023-03-08T10:00:00+05:30</published><updated>2023-03-08T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2023-03-08:/blog/identity-systems-fail-socially/</id><summary type="html">&lt;p&gt;When people talk about identity and SSO, they reach for algorithms: token lifetimes, key rotation, SAML vs OIDC, session fixation. Those details matter. In systems I have built and operated, the outages and near-misses that hurt most started earlier—as &lt;strong&gt;social and product failures&lt;/strong&gt; wearing security clothing.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/identity-systems-fail-socially.jpg" alt="Illustration of keys, badges, and people connected in a trust network" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Identity systems often …&lt;/span&gt;&lt;/figcaption&gt;&lt;/figure&gt;</summary><content type="html">&lt;p&gt;When people talk about identity and SSO, they reach for algorithms: token lifetimes, key rotation, SAML vs OIDC, session fixation. Those details matter. In systems I have built and operated, the outages and near-misses that hurt most started earlier—as &lt;strong&gt;social and product failures&lt;/strong&gt; wearing security clothing.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/identity-systems-fail-socially.jpg" alt="Illustration of keys, badges, and people connected in a trust network" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Identity systems often fail first as ownership and session-semantics problems, not as crypto bugs.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;Identity is a dependency graph of humans&lt;/h2&gt;
&lt;p&gt;An identity platform is not only a service. It is: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Who is allowed to grant access  &lt;/li&gt;
&lt;li&gt;How contractors, partners, and acquisitions show up  &lt;/li&gt;
&lt;li&gt;What "logout" means across devices and apps  &lt;/li&gt;
&lt;li&gt;Which team gets paged when login breaks on a Sunday  &lt;/li&gt;
&lt;li&gt;How quickly a leaver loses access in reality, not in policy PDFs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Crypto bugs are rare relative to &lt;strong&gt;misowned workflows&lt;/strong&gt;. The system can be textbook-correct and still fail the organization. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Failure mode 1: Ambiguous system of record for "who is this person?"&lt;/h2&gt;
&lt;p&gt;Enterprises collect identities the way rivers collect silt: HRIS, directories, partner IdPs, legacy user tables, support tools that mint exceptions. If you cannot answer "what is the canonical identifier, and who can change it?" you will eventually: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Duplicate humans  &lt;/li&gt;
&lt;li&gt;Orphan entitlements  &lt;/li&gt;
&lt;li&gt;Merge the wrong accounts  &lt;/li&gt;
&lt;li&gt;Build reconciliation jobs that become the real product&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;SSO does not fix identity entropy. It multiplies whatever model you already have. &lt;strong&gt;Design move:&lt;/strong&gt; pick a primary subject key strategy, document merge/split procedures, and make account recovery a first-class flow—not a Zendesk folklore. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Failure mode 2: Special cases without expiry&lt;/h2&gt;
&lt;p&gt;Sales needs a demo tenant. Support needs impersonation. A partner needs a long-lived integration user. Security accepts a temporary bypass. Temporary becomes permanent. Permanent becomes unmonitored. Unmonitored becomes the breach path. &lt;strong&gt;Design move:&lt;/strong&gt; every exception has an owner, an expiry, metrics, and a removal plan. Impersonation and break-glass are products with audit trails, not Slack approvals.&lt;/p&gt;
&lt;h2&gt;Failure mode 3: Logout and session mental models differ by app&lt;/h2&gt;
&lt;p&gt;Users think "I logged out." Your distributed sessions think "two refresh tokens and a cache entry are still valid." Mobile thinks something else. A partner app never got the memo. This is not merely UX. It is an access-control bug with a friendly face. &lt;strong&gt;Design move:&lt;/strong&gt; define session lifecycle as an explicit cross-app contract. Test logout like you test login. Include shared devices and support scenarios.&lt;/p&gt;
&lt;h2&gt;Failure mode 4: Rollouts that treat auth like a feature flag toy&lt;/h2&gt;
&lt;p&gt;Identity changes have asymmetric risk. A broken profile color is annoying. A broken token validation is company-wide stoppage. Yet teams still ship auth changes like UI tweaks: wide rollouts, thin dashboards, no rehearsal of rollback. &lt;strong&gt;Design move:&lt;/strong&gt; progressive exposure, synthetic login journeys per IdP, clear rollback that does not require a hero, and change freezes around known peak login events.&lt;/p&gt;
&lt;h2&gt;Failure mode 5: Ownership is "security and also platform and also the app"&lt;/h2&gt;
&lt;p&gt;When login fails, three teams page each other. When it works, nobody funds hardening. Diffused ownership produces brittle reliability. &lt;strong&gt;Design move:&lt;/strong&gt; a single operational owner for the login path, with written dependencies on IdP, DNS, email, device services, and app session layers. Security sets policy; platform runs the path; apps integrate against a stable interface. Blurry RACI is an availability risk.&lt;/p&gt;
&lt;h2&gt;What good looks like&lt;/h2&gt;
&lt;p&gt;Strong identity programs I have seen share traits: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Boring standards on the outside&lt;/strong&gt; (OIDC/SAML done plainly)  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Strict internal models&lt;/strong&gt; for subjects, credentials, devices, and grants  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Auditability as a product feature&lt;/strong&gt;  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Recovery and leaver flows tested&lt;/strong&gt;, not assumed  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Customer-visible status&lt;/strong&gt; when auth is degraded  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Load and failure testing of login&lt;/strong&gt;, not only of core APIs&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Notice how little of that is "pick the trendy token format." That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Questions before you add another identity feature&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Who is the human-level source of truth?  &lt;/li&gt;
&lt;li&gt;What is the break-glass path, and who audits it?  &lt;/li&gt;
&lt;li&gt;How does a user understand their sessions?  &lt;/li&gt;
&lt;li&gt;What happens to downstream caches on revoke?  &lt;/li&gt;
&lt;li&gt;Which team’s error budget does login reliability consume?  &lt;/li&gt;
&lt;li&gt;Can we re-run last quarter’s incidents as game days?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If those answers are soft, new federation features will add surface area, not safety. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Identity systems do fail cryptographically—and you should hire people who care about that deeply. But if you only harden tokens while leaving ownership, exceptions, session semantics, and rollout discipline vague, you will still fail. They fail socially first: unclear truth, unowned edges, temporary forever, and teams that cannot coordinate under stress. Build the social protocol as carefully as the crypto protocol. Users feel both. Attackers only need one to be weak.&lt;/p&gt;</content><category term="Blog"/><category term="identity"/><category term="security"/><category term="distributed-systems"/></entry><entry><title>What “led the web launch” taught me about constraints</title><link href="https://www.mohitranka.com/blog/web-launch-constraints/" rel="alternate"/><published>2022-07-06T10:00:00+05:30</published><updated>2022-07-06T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2022-07-06:/blog/web-launch-constraints/</id><summary type="html">&lt;p&gt;At Postman, “put the product on the web” was not a greenfield rewrite. It was a constraint problem: a desktop-native API tool used by millions of developers, enterprise pressure for browser access, browser security that blocked the old execution model, and a conference date that would not move. I was …&lt;/p&gt;</summary><content type="html">&lt;p&gt;At Postman, “put the product on the web” was not a greenfield rewrite. It was a constraint problem: a desktop-native API tool used by millions of developers, enterprise pressure for browser access, browser security that blocked the old execution model, and a conference date that would not move. I was the engineering manager accountable for cross-functional delivery—architecture choices, security and infra dependencies, product scope, and keeping the team focused when the path was still uncertain. What follows is what that launch actually taught me.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/web-launch-constraints.jpg" alt="Illustration of browser and desktop windows bridged together" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Launching a trusted desktop product on the web is a constraint problem: security, parity, and an immovable date.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;We were borrowing trust, not inventing it&lt;/h2&gt;
&lt;p&gt;Postman already had a reputation. Developers had muscle memory for collections, environments, and the desktop workflow. A weak web surface would not be judged as “v1 of a new product.” It would be judged as Postman getting worse. That constraint changed prioritization: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Core workflow parity beat architectural purity.&lt;/li&gt;
&lt;li&gt;Predictable behavior beat clever browser tricks.&lt;/li&gt;
&lt;li&gt;Explicit “not on web yet” beat silent missing features.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Trust is spendable once. We treated every launch-day gap as a brand risk, not a backlog curiosity. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;The existing system was a stakeholder&lt;/h2&gt;
&lt;p&gt;Greenfield essays assume you choose the stack. We inherited runtimes, sync assumptions, offline collaboration expectations, and a large surface area of API tooling behavior. The desktop app was not legacy to be embarrassed about—it was the system of record for how users worked. The hard question was never “can we draw a web architecture?” It was:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What must be reused so results stay correct?&lt;/li&gt;
&lt;li&gt;What must be isolated so the browser can ship?&lt;/li&gt;
&lt;li&gt;What bugs will the web amplify because usage patterns change?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Treating the existing product as a stakeholder forced interface thinking: web was a new client of a product system, not a parallel fantasy product. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Browser security forced a hybrid execution model&lt;/h2&gt;
&lt;p&gt;Desktop Postman could talk to the network like a normal app. Browsers cannot. CORS, sandboxing, and the lack of unrestricted local network access were not edge cases—they were the product. We ended up with multiple execution paths, each owning a real constraint: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Browser Agent&lt;/strong&gt; — run requests directly when the browser is allowed to.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cloud Agent&lt;/strong&gt; — execute in a cloud-hosted environment when the browser cannot reach the target cleanly (cross-origin and related limits).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Desktop Agent&lt;/strong&gt; — bridge the web UI to a local/on-prem network when the user’s world is not reachable from the public cloud.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That hybrid model was the architectural heart of the launch. It was also an organizational heart: security, infra, and product had to agree on what “send request” meant in three different trust domains. If there is one technical lesson I would keep from the project, it is this: &lt;strong&gt;when the environment cannot support your old runtime assumptions, make the execution model explicit.&lt;/strong&gt; Hiding three behaviors behind one button without a design is how you get support chaos.&lt;/p&gt;
&lt;h2&gt;Launch day is a reliability event&lt;/h2&gt;
&lt;p&gt;Postman on the Web was announced at POSTCON. Missing the date was not a soft failure mode. That does not mean we shipped fantasy scope. It means readiness was defined as: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a user journey that worked under real constraints,&lt;/li&gt;
&lt;li&gt;a rollout plan that could expand,&lt;/li&gt;
&lt;li&gt;and a team that knew what was deliberately later.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We phased capability instead of pretending the first public cut was the end state: start with constrained access patterns, then enable richer API execution, with the cloud execution path continuing to mature after the headline launch. Marketing owns the keynote. Engineering owns the degradation and expansion story. A launch is not a timestamp. It is a reliability event with an audience.&lt;/p&gt;
&lt;h2&gt;Cross-team coordination was the critical path&lt;/h2&gt;
&lt;p&gt;The longest pole was rarely a single function. Security, infrastructure, performance, and product had legitimate, conflicting optimization targets. In that environment, “the engineers will figure it out in Slack” is not a plan. What worked in practice: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Executive air cover for dedicated capacity&lt;/strong&gt; — without it, every dependency team optimizes for their prior roadmap.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A weekly cross-functional sync&lt;/strong&gt; whose job was unblocking, not status theater.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Written scope decisions&lt;/strong&gt; — what was in for conference day, what was explicit debt, who owned the follow-through.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;My calendar was part of the architecture. Ambiguity multiplies under deadline pressure; decision logs shrink it. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Performance was a product constraint, not polish&lt;/h2&gt;
&lt;p&gt;API collections can be huge. A desktop WebView habit does not automatically become a good browser experience. Large histories and large collections will punish naive rendering. We invested in boring, necessary work: more efficient history/state handling, lazy loading, virtualized UI for large lists. That work is easy to dismiss as polish until a power user loads a real workspace and the tab melts. Developer products have an unforgiving feedback loop. Users can tell when the runtime is lying, when the UI is papering over cost, and when error messages are decorative. They will also write about it publicly. That is part of the market.&lt;/p&gt;
&lt;h2&gt;What I would repeat&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Write non-goals as carefully as goals.&lt;/strong&gt; Conference-day success needs a spine, not a vision deck.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Instrument journeys, not only services.&lt;/strong&gt; “Request failed” is incomplete without &lt;em&gt;which agent path&lt;/em&gt; and &lt;em&gt;which constraint&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rehearse partial failure&lt;/strong&gt; — auth issues, agent unavailability, dependency brownouts—not only happy-path demos.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Staff the week after launch like it is part of launch.&lt;/strong&gt; The real traffic pattern arrives after the keynote.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Protect engineers from thrash&lt;/strong&gt; by batching stakeholder input; panic multiplies bad architectural shortcuts.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;What I would avoid&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Betting the public launch on an unfinished platform rewrite that is “almost ready.”&lt;/li&gt;
&lt;li&gt;Hiding scope cuts inside the word “polish.”&lt;/li&gt;
&lt;li&gt;Success metrics only a marketing team can love.&lt;/li&gt;
&lt;li&gt;Hero culture that makes the second week impossible to staff.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Impact, carefully stated&lt;/h2&gt;
&lt;p&gt;We hit the conference launch window. The web surface became a real product path, not a demo—with cloud execution continuing to land on its own schedule. Adoption afterward made the strategic point obvious: users wanted Postman without installing a desktop app first, and the company was no longer only a local-first tool. Exact figures belong in contexts where they can be sourced and defended. The leadership lesson does not depend on a screenshot of a dashboard: &lt;strong&gt;the hybrid execution model plus phased delivery was the only way to respect browser constraints without abandoning the desktop product’s trust.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;“Led the web launch” sounds like a milestone. The work was constraint management: borrowed trust, inherited systems, browser security, conference time, cross-team conflict, and performance under real collections. Code expressed the answers. The answers were the constraints we were willing to name early—and the execution model we built so users did not have to understand them all at once.&lt;/p&gt;</content><category term="Blog"/><category term="engineering-leadership"/><category term="product"/><category term="developer-tooling"/></entry><entry><title>Incubating 0→1 beside a mature product</title><link href="https://www.mohitranka.com/blog/postman-labs-0-to-1/" rel="alternate"/><published>2021-06-15T10:00:00+05:30</published><updated>2021-06-15T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2021-06-15:/blog/postman-labs-0-to-1/</id><summary type="html">&lt;p&gt;When I was at Postman, the core product was already the default API client for a huge HTTP/HTTPS world—on the order of tens of millions of developers on desktop. That success created a sharp problem: &lt;strong&gt;how do you explore what comes after HTTP without slowing the product everyone …&lt;/strong&gt;&lt;/p&gt;</summary><content type="html">&lt;p&gt;When I was at Postman, the core product was already the default API client for a huge HTTP/HTTPS world—on the order of tens of millions of developers on desktop. That success created a sharp problem: &lt;strong&gt;how do you explore what comes after HTTP without slowing the product everyone already depends on?&lt;/strong&gt; Postman Labs was our answer: a small unit with a charter to incubate 0→1 work—new protocols and paradigms—without turning every experiment into a core-roadmap hostage situation.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/postman-labs-0-to-1.jpg" alt="Illustration of a small lab greenhouse beside a solid product building" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;0→1 incubation works when it sits beside a mature product with a real graduation path—not as endless side quests.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;The challenge: the market moved past “REST in a GUI”&lt;/h2&gt;
&lt;p&gt;Developers were increasingly living with: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;WebSockets&lt;/strong&gt; for persistent, bidirectional sessions  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;gRPC&lt;/strong&gt; for efficient, schema-driven service APIs  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GraphQL&lt;/strong&gt; and other non-CRUD shapes of interface&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Internally we had ideas. What we lacked was a structured way to &lt;strong&gt;validate, build, and kill or graduate&lt;/strong&gt; them while the core team stayed focused on the desktop product’s quality and scale. Putting every bet into the main feature factory would have meant either: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;starving core reliability and UX, or  &lt;/li&gt;
&lt;li&gt;shipping “innovation” at the speed of a mature backlog.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Labs existed to refuse that false choice. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Structure for speed (and containment)&lt;/h2&gt;
&lt;p&gt;We did not invent another feature team with the same process tax as core. Labs was intentionally different: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Smaller team&lt;/strong&gt; drawn from engineers who already knew Postman’s product DNA.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Less process theater&lt;/strong&gt; — lean experiments instead of full SDLC cosplay for every spike.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One North Star metric per initiative&lt;/strong&gt; — a single definition of success so debates stayed grounded.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explicit separation&lt;/strong&gt; from core delivery so a failed experiment did not become a multi-quarter core commitment by accident.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Charter in two axes: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Breadth&lt;/strong&gt; — protocols and shapes beyond HTTP.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Depth&lt;/strong&gt; — personas and workflows (testing, automation, CI-shaped use) that the HTTP client alone did not own.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Independence was not isolation from users. It was isolation from the wrong kind of backlog pressure. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;A three-phase model that de-risked 0→1&lt;/h2&gt;
&lt;p&gt;Every Labs initiative had to earn the next phase: ### Phase 1 — Feasibility and market validation
Customer conversations, lightweight proofs, honest “who hurts without this?” If the answer was only “it would be cool,” it did not proceed. ### Phase 2 — MVP and dogfooding
Internal builds used by Postman’s own engineers. Usability and correctness issues showed up before a public audience. ### Phase 3 — Limited beta and public validation
Constrained rollout, measure adoption and engagement, then decide: graduate into the main product, iterate, or stop. That sequence sounds obvious. The discipline is stopping between phases. Core roadmaps often skip phase 1 and call a half-built feature a launch.&lt;/p&gt;
&lt;h2&gt;WebSockets: persistent sessions are not “requests with extra steps”&lt;/h2&gt;
&lt;p&gt;WebSockets were an early Labs bet because real-time systems (chat, feeds, IoT-style control planes, live tooling) do not fit the request/response mental model Postman had optimized for. &lt;strong&gt;Hard parts:&lt;/strong&gt; That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Long-lived bidirectional connections instead of discrete calls  &lt;/li&gt;
&lt;li&gt;Auth flows that must hold for a session, not only a single hit  &lt;/li&gt;
&lt;li&gt;UX for streams of events over time, not one response panel&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;What we built toward:&lt;/strong&gt; a first-class WebSocket client experience—connect, send/receive, inspect event history, support practical auth patterns. &lt;strong&gt;Outcome:&lt;/strong&gt; WebSockets did not stay a lab toy; support graduated into Postman’s broader API development surface. That graduation path was the point of Labs. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;gRPC: schemas and streams in a product trained on text HTTP&lt;/h2&gt;
&lt;p&gt;gRPC was growing fast in backend-heavy environments. Postman’s muscle memory was text-centric HTTP. gRPC forced different questions: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How do users work with &lt;strong&gt;Protobuf&lt;/strong&gt; contracts inside the product?  &lt;/li&gt;
&lt;li&gt;How do unary and streaming calls show up in a composer UX?  &lt;/li&gt;
&lt;li&gt;How do serialization mistakes become debuggable instead of opaque binary pain?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Execution themes:&lt;/strong&gt; schema-aware composition, streaming-aware request/response handling, and making “what did I just send?” inspectable for developers who live in Postman daily. &lt;strong&gt;Outcome:&lt;/strong&gt; gRPC testing/debugging became a real product capability in the same family as REST workflows—not a separate science project users had to leave Postman for.&lt;/p&gt;
&lt;h2&gt;What scaled beyond the first bets&lt;/h2&gt;
&lt;p&gt;When early graduates worked, Labs stopped being only a temporary squad. The incubation pattern—validate, dogfood, beta, graduate—became a reusable company muscle. Later product bets (including automation-shaped work such as Flows-class ideas) could reuse the same organizational shape: &lt;strong&gt;explore beside core, then merge what earns users.&lt;/strong&gt; Eventually Labs-shaped work attracted clearer funding and leadership attention. That is the healthy end state: not a permanent rebel base, but a proven path for 0→1 inside a company that also has to protect a mature product.&lt;/p&gt;
&lt;h2&gt;What I would repeat as an EM&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Separate exploration capacity from core SLA capacity&lt;/strong&gt; — or core always wins and innovation becomes slideware.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One success metric per bet&lt;/strong&gt; — multi-metric dashboards hide kill decisions.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dogfood before marketing&lt;/strong&gt; — especially for developer tools; your engineers are harsh, useful users.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Graduation criteria in writing&lt;/strong&gt; — “done in Labs” must mean something operationally.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Protect the team from identity crisis&lt;/strong&gt; — Labs is not “the people who do side quests”; it is a product strategy role.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;What I would avoid&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Innovation theater with no kill switch  &lt;/li&gt;
&lt;li&gt;Hiding Labs work so core is surprised at graduation  &lt;/li&gt;
&lt;li&gt;Measuring success only by launches, not by retained use  &lt;/li&gt;
&lt;li&gt;Staffing Labs only with people core “can spare” forever&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Incubating 0→1 beside a mature product is a leadership design problem: process, incentives, and graduation rules—not only prototype velocity. Postman Labs worked when it had a &lt;strong&gt;narrow charter&lt;/strong&gt;, &lt;strong&gt;phased validation&lt;/strong&gt;, and a path for WebSockets, gRPC, and similar bets to become real product surfaces without forcing the entire company to pretend it was still a startup with nothing to lose. If your core product is already loved, that is not a reason to stop exploring. It is a reason to explore &lt;strong&gt;on purpose&lt;/strong&gt;.&lt;/p&gt;</content><category term="Blog"/><category term="engineering-leadership"/><category term="product"/><category term="developer-tooling"/></entry><entry><title>How I review a distributed design in 45 minutes</title><link href="https://www.mohitranka.com/blog/how-i-review-a-distributed-design/" rel="alternate"/><published>2021-03-03T10:00:00+05:30</published><updated>2021-03-03T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2021-03-03:/blog/how-i-review-a-distributed-design/</id><summary type="html">&lt;p&gt;The most expensive design reviews I have run were not missing a box on a diagram. They were missing a decision. One of them froze delivery on an EDP self-serve portal for about a month while two strong engineers disagreed in public about monolith versus microservices—and ownership quietly collapsed …&lt;/p&gt;</summary><content type="html">&lt;p&gt;The most expensive design reviews I have run were not missing a box on a diagram. They were missing a decision. One of them froze delivery on an EDP self-serve portal for about a month while two strong engineers disagreed in public about monolith versus microservices—and ownership quietly collapsed. This is how I run a distributed design review when time is short, using that conflict as the worked example. The forty-five minutes are a filter for &lt;strong&gt;danger and indecision&lt;/strong&gt;, not a substitute for deep design.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/how-i-review-a-distributed-design.jpg" alt="Illustration of an architecture whiteboard with boxes, arrows, and coffee cups" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;A short design review is a filter for clear promises, truth models, and decisions—not a theater of diagrams.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;The situation the review had to unstick&lt;/h2&gt;
&lt;p&gt;We needed a self-serve portal on top of the Enterprise Data Platform: dataset registration, governance controls, access workflows, metadata—so teams could manage lifecycle without filing tickets into oblivion. Two credible positions formed: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Stay close to the monolithic EDP backend&lt;/strong&gt; — simpler integration, less duplication, faster delivery on shared infra.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Split into microservices early&lt;/strong&gt; — independent iteration on portal features without waiting on the monolith.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Both sides had technical merit. The failure mode was social: critique moved into open forums in a way that undermined the owner, the owner disengaged, and execution stopped. Stakeholders saw silence. That is a design-process failure, not only an architecture debate. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Minutes 0–5: What user promise are we keeping?&lt;/h2&gt;
&lt;p&gt;Before hexagons, I want one paragraph: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Who uses the portal?&lt;/li&gt;
&lt;li&gt;What can they do without a human intermediary?&lt;/li&gt;
&lt;li&gt;What is explicitly out of scope for v1?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For us: GTM/data producers and consumers managing dataset lifecycle—registration, access, metadata—not “rebuild EDP as microservices.” If the promise is fuzzy, stop. Architecture will invent scope. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Minutes 5–15: Where does truth live?&lt;/h2&gt;
&lt;p&gt;I care about systems of record more than service count. Questions that mattered on the portal: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Which actions must be consistent with core EDP backend behavior on day one?&lt;/li&gt;
&lt;li&gt;Which workflows change weekly and need independent deploy cadence?&lt;/li&gt;
&lt;li&gt;Where does dataset metadata live so the portal is not a second brain?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We eventually used a hybrid truth model: &lt;strong&gt;core platform capabilities stayed integrated with the existing EDP backend&lt;/strong&gt;; &lt;strong&gt;more dynamic workflows&lt;/strong&gt; (access requests, tagging-style features) could stand as separate services; &lt;strong&gt;metadata&lt;/strong&gt; was centralized with a system fit for dataset lifecycle (in our case, Apache DataHub) so the portal was not inventing yet another catalog. Red flags in any review:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Both systems will stay in sync” with no mechanism.&lt;/li&gt;
&lt;li&gt;Every feature forced into one deployability story.&lt;/li&gt;
&lt;li&gt;Metadata treated as a UI detail.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Minutes 15–25: What fails—technically and organizationally?&lt;/h2&gt;
&lt;p&gt;Classic distributed questions still apply: timeouts, dual writes, partial deploy, replay. On this project the binding failure was different: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What happens if the owning engineer stops driving?&lt;/li&gt;
&lt;li&gt;What happens if disagreement becomes a public referendum every week?&lt;/li&gt;
&lt;li&gt;What happens if leadership hears only one side’s framing?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A design review that ignores ownership and decision rights will produce a beautiful diagram and a still project. I schedule the technical argument &lt;strong&gt;inside a structured forum&lt;/strong&gt; with criteria—not in drive-by threads. If you need blame-free space, create it deliberately. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Minutes 25–35: How will we choose without infinite debate?&lt;/h2&gt;
&lt;p&gt;Opinion without evidence burns weeks. The intervention that worked: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Write evaluation criteria&lt;/strong&gt; before picking a winner: scalability, maintainability, delivery speed, long-term ownership.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Force small proofs&lt;/strong&gt; — both approaches get a time-boxed spike/POC against the criteria.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bring a neutral senior engineer&lt;/strong&gt; into the room to pressure-test both sides without owning either ego.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Time-box the decision&lt;/strong&gt; — the review ends with a path, not a sequel meeting.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is operability of the &lt;em&gt;decision&lt;/em&gt;, not only of the service. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Minutes 35–40: People side (do not skip)&lt;/h2&gt;
&lt;p&gt;Distributed design is done by humans with status and career goals. In parallel with the technical path: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Rebuild ownership with the engineer who had stepped back—silence is not an acceptable escalation strategy.&lt;/li&gt;
&lt;li&gt;Coach the critic on influence: staff-level impact includes &lt;em&gt;how&lt;/em&gt; you challenge, not only that you are right.&lt;/li&gt;
&lt;li&gt;Make expectations explicit in 1:1s so the project is not a proxy war.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Skip this and the hybrid architecture will still die in the next disagreement. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Minutes 40–45: Decision and conditions&lt;/h2&gt;
&lt;p&gt;We did not pick a pure monolith or a pure microservice rewrite. We picked a &lt;strong&gt;hybrid&lt;/strong&gt;: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Core platform features (registration, governance controls tightly bound to EDP) stayed where integration cost dominated.&lt;/li&gt;
&lt;li&gt;Dynamic portal features that needed independent iteration moved toward separate services.&lt;/li&gt;
&lt;li&gt;Metadata lifecycle was centralized rather than re-implemented.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Approve with conditions, in writing: what is in v1, what is explicitly later, who owns the seams, when the next review is if assumptions fail. Verbal “sounds good” evaporates. Written conditions survive contact with calendars. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;The diagram I want if I only get one&lt;/h2&gt;
&lt;p&gt;Trade three layered architecture posters for either: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;sequence&lt;/strong&gt; of one user action through registration → metadata → access, or  &lt;/li&gt;
&lt;li&gt;a &lt;strong&gt;side-by-side POC scorecard&lt;/strong&gt; against the agreed criteria.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Sequence diagrams and scorecards reveal lies that box diagrams hide. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Anti-patterns this freeze taught me&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Public architecture criticism that bypasses the owner.&lt;/li&gt;
&lt;li&gt;Binary holy wars (monolith vs microservices) without workload specifics.&lt;/li&gt;
&lt;li&gt;Design review as spectator sport for leadership without a decision owner.&lt;/li&gt;
&lt;li&gt;EM waiting too long to facilitate because “they’re seniors, they’ll figure it out.”&lt;/li&gt;
&lt;li&gt;Approving to end discomfort rather than risk.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;What unblocked looked like&lt;/h2&gt;
&lt;p&gt;Once criteria, POCs, and a hybrid decision landed, the freeze broke quickly—on the order of a week to resume real execution—and the portal shipped on a timeline measured in months with meaningful adoption. The architecture mattered. The restored ownership mattered more. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;In forty-five minutes you will not finish a distributed design. You can learn whether the team has a clear promise, a coherent truth model, a way to decide, and a human ownership path. On platform work, the last item is not soft. It is how delivery fails first. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;</content><category term="Blog"/><category term="distributed-systems"/><category term="engineering-leadership"/><category term="architecture"/></entry><entry><title>Introducing on-call without burning the team</title><link href="https://www.mohitranka.com/blog/introducing-on-call-without-burnout/" rel="alternate"/><published>2020-11-10T10:00:00+05:30</published><updated>2020-11-10T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2020-11-10:/blog/introducing-on-call-without-burnout/</id><summary type="html">&lt;p&gt;At Postman, my team owned a large-scale platform surface used by millions of developers. What we did not own—formally—was a &lt;strong&gt;predictable operational response&lt;/strong&gt;. Production issues and public GitHub noise were handled ad hoc. Someone jumped in, or everyone hesitated. Retrospectives lacked a clear accountable role. The system worked …&lt;/p&gt;</summary><content type="html">&lt;p&gt;At Postman, my team owned a large-scale platform surface used by millions of developers. What we did not own—formally—was a &lt;strong&gt;predictable operational response&lt;/strong&gt;. Production issues and public GitHub noise were handled ad hoc. Someone jumped in, or everyone hesitated. Retrospectives lacked a clear accountable role. The system worked until it did not, and then it worked by heroics. We needed on-call. We also needed engineers who still wanted to build product after the rotation.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/introducing-on-call-without-burnout.jpg" alt="Illustration of a calm night operations desk with a pager and schedule board" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;On-call is a product and staffing design: clear ownership without making heroics the default.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;What was broken before process&lt;/h2&gt;
&lt;p&gt;Without a rotation, four things piled up: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Unclear ownership&lt;/strong&gt; — incidents waited on “who feels responsible today.”  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context-switch tax&lt;/strong&gt; — feature work and firefighting shared the same brains without boundaries.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Uneven load&lt;/strong&gt; — the same people always raised their hands.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Weak learning loops&lt;/strong&gt; — retros had symptoms, not a role that carried fixes week to week.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For a developer-facing platform, user-visible breakage is not a side channel. It is the product. Treating ops as optional was a product decision, whether we admitted it or not. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Design goals&lt;/h2&gt;
&lt;p&gt;I wanted a system that was: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Explicit&lt;/strong&gt; — someone is primary, always.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fair&lt;/strong&gt; — load rotates; it does not stick to volunteers.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bounded&lt;/strong&gt; — on-call is a job for a window, not a personality type.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Educational&lt;/strong&gt; — the whole team sees production, not only a martyr subset.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Humane&lt;/strong&gt; — no permanent page-from-bed culture dressed up as commitment.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Reliability that depends on burnout is just deferred attrition. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;The model we ran&lt;/h2&gt;
&lt;h3&gt;Primary and secondary&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Primary&lt;/strong&gt; — dedicated to monitoring, incident response, and triage (including GitHub-facing noise). Feature delivery expectations drop for that window on purpose.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Secondary&lt;/strong&gt; — backup for major incidents; not a stealth second primary for every ping.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If primary is still expected to hit the same sprint commitments, you do not have on-call. You have theater. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h3&gt;Two-week shifts with a handoff pattern&lt;/h3&gt;
&lt;p&gt;Engineers rotated on a &lt;strong&gt;two-week&lt;/strong&gt; cadence. A common pattern was primary one window, then secondary the next—so knowledge transferred and no one lived forever in the blast radius. Back-to-back primary stretches were treated as a smell, not a badge. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h3&gt;Handoff as a team ritual&lt;/h3&gt;
&lt;p&gt;Weekly team time included an &lt;strong&gt;on-call handoff&lt;/strong&gt;, not only standup status. Primary walked through: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Incidents and resolutions  &lt;/li&gt;
&lt;li&gt;Adjacent system issues that might hit us next  &lt;/li&gt;
&lt;li&gt;Notable GitHub / support themes  &lt;/li&gt;
&lt;li&gt;Follow-ups that needed owners beyond the shift&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That ritual turned private pager pain into shared product knowledge. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h3&gt;Retros and playbooks&lt;/h3&gt;
&lt;p&gt;Major incidents got structured retros. Recurring issues earned &lt;strong&gt;playbooks&lt;/strong&gt;—step-by-step paths so the next primary was not rediscovering folklore at 1 a.m. Playbooks are how on-call becomes a team asset instead of tribal knowledge in one engineer’s head. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h3&gt;Psychological safety and load management&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;No expectation of endless consecutive primaries.  &lt;/li&gt;
&lt;li&gt;Balance with feature work across the quarter so people are not typed as “ops only.”  &lt;/li&gt;
&lt;li&gt;Managers (me included) treated page load and fairness as staffing concerns, not only engineer grit.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;What improved&lt;/h2&gt;
&lt;p&gt;With clear ownership: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Faster response&lt;/strong&gt; — we saw on the order of a &lt;strong&gt;~40% reduction in incident response time&lt;/strong&gt; once roles were unambiguous (directionally; treat it as an order-of-magnitude win from process, not a lab result).  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;More systematic triage&lt;/strong&gt; — fewer “is anyone looking at this?” gaps.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Better morale&lt;/strong&gt; — predictable pain beats random pain.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Broader operational skill&lt;/strong&gt; — more engineers touched production reality.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Earlier fixes&lt;/strong&gt; — primaries had space to chip at known sharp edges before they became SEVs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Stability improved because response became a designed system, not a personality contest. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;What I would tell another EM before day one&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Write who is primary in a place the team actually looks.&lt;/strong&gt;  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cut feature load for primary&lt;/strong&gt; or you will train people to ignore the pager.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ship handoff and playbooks in the first month&lt;/strong&gt;, not after the third outage.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Measure response and fairness&lt;/strong&gt;, not only uptime.  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Defend the rotation against “just this once” exceptions&lt;/strong&gt; from leadership—exceptions are how volunteers reappear.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Failure modes to avoid&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;On-call as punishment for the least political engineers  &lt;/li&gt;
&lt;li&gt;Secondary as free extra primary  &lt;/li&gt;
&lt;li&gt;Retros without owners or due dates  &lt;/li&gt;
&lt;li&gt;Alert noise so high that everyone mutes everything  &lt;/li&gt;
&lt;li&gt;Celebrating heroes instead of fixing the systems that required them&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;Introducing on-call is not a tooling purchase. It is a &lt;strong&gt;product and staffing decision&lt;/strong&gt;: user trust requires an accountable human path, and that path must be sustainable. At Postman, primary/secondary roles, two-week rotations, handoffs, retros, and playbooks turned operational ownership from ad hoc heroics into something the team could carry—and still ship. If your platform is already large and your response is still “whoever notices,” you do not have a reliability gap only. You have a leadership design gap. Close it on purpose.&lt;/p&gt;</content><category term="Blog"/><category term="reliability"/><category term="engineering-leadership"/><category term="platforms"/></entry><entry><title>GTM datasets need data contracts</title><link href="https://www.mohitranka.com/blog/gtm-datasets-need-data-contracts/" rel="alternate"/><published>2020-07-01T10:00:00+05:30</published><updated>2020-07-01T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2020-07-01:/blog/gtm-datasets-need-data-contracts/</id><summary type="html">&lt;p&gt;When a go-to-market dashboard is wrong, nobody says “the warehouse is eventually consistent.” They say the number is wrong—and they stop trusting the platform. On LinkedIn’s Enterprise Data Platform (EDP) work, the failure mode was rarely a missing chart type. It was &lt;strong&gt;informal truth&lt;/strong&gt;: datasets without clear producers …&lt;/p&gt;</summary><content type="html">&lt;p&gt;When a go-to-market dashboard is wrong, nobody says “the warehouse is eventually consistent.” They say the number is wrong—and they stop trusting the platform. On LinkedIn’s Enterprise Data Platform (EDP) work, the failure mode was rarely a missing chart type. It was &lt;strong&gt;informal truth&lt;/strong&gt;: datasets without clear producers, consumers, freshness expectations, or a path that BI could rely on when Hadoop-era pipelines still “worked.” That is a data-contract problem, whether or not you use the word contract.&lt;/p&gt;
&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/gtm-datasets-need-data-contracts.jpg" alt="Illustration of two parties exchanging a contract over data folders" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;GTM datasets become trustworthy when producers and consumers share explicit contracts—not informal folklore.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;!--more--&gt;

&lt;h2&gt;The contract is the product boundary&lt;/h2&gt;
&lt;p&gt;A data contract is a written agreement between people who produce a dataset and people who depend on it: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What the dataset means (grain, keys, critical fields)&lt;/li&gt;
&lt;li&gt;Who owns changes&lt;/li&gt;
&lt;li&gt;How fresh and complete it must be for its main consumers&lt;/li&gt;
&lt;li&gt;Who may use it, and for what&lt;/li&gt;
&lt;li&gt;What happens when the shape changes&lt;/li&gt;
&lt;li&gt;Which path is the system of record when two feeds disagree&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Without that, you have tables and jobs—not a product interface. EDP’s strategic job was to become the governed center for GTM data. BI teams on Power BI and Tableau did not move because a platform existed; they moved when the &lt;strong&gt;interface of getting trustworthy data&lt;/strong&gt; became clearer, cheaper, and eventually mandatory as legacy paths aged out.&lt;/p&gt;
&lt;h2&gt;Which GTM datasets need contracts (almost all that matter)&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Definitely:&lt;/strong&gt; That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Pipeline and revenue metrics that show up in leadership reviews  &lt;/li&gt;
&lt;li&gt;Account, lead, and opportunity-shaped datasets used across tools  &lt;/li&gt;
&lt;li&gt;Any feed BI materializes into workbooks that drive weekly motions  &lt;/li&gt;
&lt;li&gt;Datasets used for access decisions, eligibility, or customer-facing ops  &lt;/li&gt;
&lt;li&gt;Shared “golden” entities multiple teams join in different ways&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Lighter-weight is fine for:&lt;/strong&gt; That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Truly exploratory sandboxes with no production consumers  &lt;/li&gt;
&lt;li&gt;One-off extracts with an explicit expiry&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If a number can start an argument in a QBR, it deserves a contract. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;What we needed in practice (not a 40-page template)&lt;/h2&gt;
&lt;p&gt;Keep contracts short enough that producers and BI partners will read them: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Dataset name&lt;/strong&gt; and owning team  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grain&lt;/strong&gt; (what one row means)  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Critical fields&lt;/strong&gt; and allowed null behavior  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Primary consumers&lt;/strong&gt; (e.g. Power BI / Tableau paths, sales analytics)  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Freshness / readiness expectation&lt;/strong&gt; — even if it starts as “available on EDP before legacy deprecation,” not a perfect percentile  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Change policy&lt;/strong&gt; — notice, versioning, who approves breaking changes  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;System of record&lt;/strong&gt; during dual-run periods  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Support path&lt;/strong&gt; — where breakages go (not a random Slack thread)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;On EDP, connectors, migration staffing, and deprecation dates were how contracts became real. A wiki table alone does not change incentives. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Why platform adoption without contracts fails&lt;/h2&gt;
&lt;p&gt;BI’s rational objection was: “We can already get the data.” Informal sources always feel free until: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Two dashboards disagree  &lt;/li&gt;
&lt;li&gt;A legacy pipeline is turned off  &lt;/li&gt;
&lt;li&gt;A field changes meaning and nobody tells the workbook owner  &lt;/li&gt;
&lt;li&gt;Query performance after cutover becomes “the platform’s problem” with no owner&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Contracts force those conversations &lt;strong&gt;before&lt;/strong&gt; the incident. They also make deprecation fair: you cannot retire a path nobody documented as non-authoritative. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;Evaluation is part of the contract&lt;/h2&gt;
&lt;p&gt;Do not separate “data quality” from “dashboard quality.” Ship and migration gates should include: That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Consumer path checks (does the BI workflow still resolve?)  &lt;/li&gt;
&lt;li&gt;Row-count / null-rate sanity on critical fields  &lt;/li&gt;
&lt;li&gt;Explicit dual-run comparisons while both paths live  &lt;/li&gt;
&lt;li&gt;A named human who can freeze a bad publish&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When the contract breaks, something visible should fail before the QBR does. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Organizational pattern that worked&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Producers&lt;/strong&gt; own correctness and change communication  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Platform (EDP)&lt;/strong&gt; owns enforcement, discovery, access patterns, and migration leverage  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;BI / analytics partners&lt;/strong&gt; own consumer semantics and workbook impact  &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Leadership&lt;/strong&gt; owns deprecation as continuity policy, not a style preference&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Shared Slack channels are not a substitute for ownership. Funded migration and executive sponsorship were how EDP contracts left the slide deck. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams.&lt;/p&gt;
&lt;h2&gt;A sequence I recommend for new GTM datasets&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Write the decision the dataset supports in one sentence.  &lt;/li&gt;
&lt;li&gt;Name the first production consumer (often a BI path).  &lt;/li&gt;
&lt;li&gt;Draft the contract &lt;em&gt;before&lt;/em&gt; scaling access.  &lt;/li&gt;
&lt;li&gt;Put the dataset on the governed platform path.  &lt;/li&gt;
&lt;li&gt;Dual-run only with an end date.  &lt;/li&gt;
&lt;li&gt;Turn off the informal path.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Teams love to start at “expose the table.” Steps 1–3 are where trust is designed. That detail matters in practice because the surrounding system, incentives, and failure modes usually determine whether the idea survives contact with production and with other teams. I treat this as an operating constraint rather than a slogan: if you cannot explain how it shows up in ownership, metrics, and day-to-day decisions, it will not survive the next roadmap fight.&lt;/p&gt;
&lt;h2&gt;Closing&lt;/h2&gt;
&lt;p&gt;EDP did not earn “source of truth” status by existing. It earned it when GTM consumers—especially BI—could depend on &lt;strong&gt;named datasets with owners, expectations, and an end to competing pipelines&lt;/strong&gt;. Call that a data contract, a product interface, or an operating promise. Just do not ship GTM data as folklore and hope governance appears later.&lt;/p&gt;</content><category term="Blog"/><category term="data-platforms"/><category term="platforms"/><category term="product"/></entry><entry><title>RDBMS vs. NOSQL?</title><link href="https://www.mohitranka.com/blog/rdbms-vs-nosql/" rel="alternate"/><published>2013-06-29T00:16:00+05:30</published><updated>2013-06-29T00:16:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2013-06-29:/blog/rdbms-vs-nosql/</id><summary type="html">&lt;blockquote&gt;&lt;p&gt;From our own experience designing and operating a highly available, highly scalable ecommerce platform, we have come to realize that relational databases should only be used when an application really needs the complex query, table join and transaction capabilities of a full-blown relational database. In all other cases, when such …&lt;/p&gt;&lt;/blockquote&gt;</summary><content type="html">&lt;blockquote&gt;&lt;p&gt;From our own experience designing and operating a highly available, highly scalable ecommerce platform, we have come to realize that relational databases should only be used when an application really needs the complex query, table join and transaction capabilities of a full-blown relational database. In all other cases, when such relational features are not needed, a NoSQL database service like DynamoDB offers a simpler, more available, more scalable and ultimately a lower cost solution.&lt;/p&gt;&lt;/blockquote&gt;

&lt;pre&gt;&lt;code&gt;            — Werner Vogels, CTO Amazon.com on when to use RDBMS
&lt;/code&gt;&lt;/pre&gt;

&lt;figure class="post-figure"&gt;
  &lt;img src="https://www.mohitranka.com/images/blog/rdbms-vs-nosql.jpg" alt="Illustration comparing ordered relational tables with flexible document nodes" loading="lazy" width="1200" height="675"&gt;
  &lt;figcaption&gt;
    &lt;span class="fig-caption"&gt;Datastore choice is a tradeoff about access patterns and operations—not a fashion contest.&lt;/span&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p&gt;I came across this quote in &lt;a href="http://www.allthingsdistributed.com/2013/03/dynamodb-one-year-later.html"&gt;an article&lt;/a&gt; while researching DynamoDB. As much as I respect Werner, I would suggest take his recommendation on the criteria on database selection with a pinch of salt, due to the obvious conflict of interest — He has a &lt;a href="http://aws.amazon.com/dynamodb/"&gt;database&lt;/a&gt; to sell.&lt;/p&gt;

&lt;!--more--&gt;

&lt;p&gt;Most products do not need the &lt;em&gt;scale&lt;/em&gt;, which cannot be served by relational databases. Relational databases have survived more than 20 years and still doing well, even at &lt;a href="https://www.facebook.com/MySQLatFacebook"&gt;large scale&lt;/a&gt;! There are &lt;a href="http://dev.mysql.com/"&gt;extremely&lt;/a&gt; &lt;a href="http://www.postgresql.org/"&gt;mature&lt;/a&gt; open source rdbms, requirement independent schemas, optimized queries for filtering, aggregation and list of data, great communities, acquired knowledge, default integration with programming frameworks and lots of tools developed.&lt;/p&gt;

&lt;h2&gt;RDBMS are great-for-all, unless&amp;hellip;&lt;/h2&gt;

&lt;p&gt;Relational systems are super easy to work with, have great (free) tools for monitoring, backup, operations available, work well for most of the use cases, there is a lot of help (community, web resources, consulting) around and best of all.. you already know it well. &lt;em&gt;Unless&lt;/em&gt; I have the following requirements I would use a single node relational database.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;h3&gt;Super high availability&lt;/h3&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Relational database&amp;rsquo;s server process is a single point of failure, which means your system &lt;em&gt;will&lt;/em&gt; go down, when the process crashes or taken down for scheduled or unscheduled maintainence. In short, prefer the boring system until a concrete requirement forces a different shape—and document that requirement so the next person does not re-litigate it from first principles.&lt;/p&gt;

&lt;p&gt;If you are building a system which needs to have close to 100% availability, look into clustered relational databases or distributed databases which choose availability over consistency, eg. cassandra. In short, prefer the boring system until a concrete requirement forces a different shape—and document that requirement so the next person does not re-litigate it from first principles.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;h3&gt;Flexible schema&lt;/h3&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One of relational databases&amp;#8217; great strength is its schema, data model and &lt;a href="https://en.wikipedia.org/wiki/Database_normalization"&gt;related concepts&lt;/a&gt;, which allow for schema modelling, without knowing the query patterns in advance, and works well. However, &lt;a href="http://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80%93value_model"&gt;not all data can be effectively represented in relational data model&lt;/a&gt;. Even though there is some support for &lt;a href="http://www.postgresql.org/docs/9.0/static/hstore.html"&gt;flexible&lt;/a&gt; &lt;a href="http://www.postgresql.org/docs/current/static/functions-json.html"&gt;schema&lt;/a&gt; datastructure, relational databases are efficitent on fixed schema which can confirm to normalization rules for a tradeoff between consistency and performance. Flexible schema &lt;a href="http://karwin.blogspot.in/2009/05/eav-fail.html"&gt;do&lt;/a&gt; &lt;a href="http://tonyandrews.blogspot.in/2004/10/otlt-and-eav-two-big-design-mistakes.html"&gt;not&lt;/a&gt; &lt;a href="https://www.simple-talk.com/opinion/opinion-pieces/bad-carma/"&gt;work well&lt;/a&gt; with relational databases.&lt;/p&gt;

&lt;p&gt;If you &lt;em&gt;need&lt;/em&gt; (do you, really?) flexible schema, look at other nosql options. In short, prefer the boring system until a concrete requirement forces a different shape—and document that requirement so the next person does not re-litigate it from first principles. In short, prefer the boring system until a concrete requirement forces a different shape—and document that requirement so the next person does not re-litigate it from first principles.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;h3&gt;Horizontal scaling&lt;/h3&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are looking to build a system that should scale horizontally as your data needs grow, a single-node relational database will not suffice forever. There is only so much room to grow, so much RAM and CPU to throw on a single machine. Sooner or later, you will hit that single server machine limit. At that point, you will have to either choose a NoSQL solution or a relational database system cluster.&lt;/p&gt;

&lt;h1&gt;Epilogue&lt;/h1&gt;

&lt;p&gt;NoSQL is no panacea and RDBMS are still the best choice for most of the applications. Unless you have strong need to &lt;em&gt;not&lt;/em&gt; used a relational database, stick to it. In short, prefer the boring system until a concrete requirement forces a different shape—and document that requirement so the next person does not re-litigate it from first principles.&lt;/p&gt;</content><category term="Blog"/><category term="data-platforms"/><category term="databases"/><category term="architecture"/></entry></feed>