← Writing

Recurring Decisions

This page catalogues a series of decisions you will want to make in a product team, spanning the why, what, and how. To make those decisions I've given a set of possible signals you can observe. These are exactly the same signals as in the field guide, but organised under the decisions, and repeated if necessary. Which ones you pay attention to will be a function of history, emphasis, and ability to collect. The signals are ordered from early to late, and it is important to keep an eye on signals from different time horizons.

How to read a row
typical latencythe signal (dot colour = track, fill = how much to trust it, ◆ = human-run)
300msIDE inline feedback
highpush
latency rangetrustpush interrupts you, pull you check
300mstypical latency
dot colour = track, fill = how much to trust it
human-run
hightrust in the signal (high, med, low)
pushpush interrupts you, pull you check
Your stage

Build

Can we make the thing?

Is this change technically correct, and safe to merge?

Extremely important for AI coding agents to apply many of the deterministic and automatic signals here. Each of the earlier signals you incorporate into your development lifecycle, if you apply them correctly, will tend to save you bugs and issues in the longer term.

  1. 300msIDE inline feedback
    highpush
  2. 500msSyntax / parser
    highpush
  3. 2sType system
    highpush
  4. 3sHot reload / REPL
    medpush
  5. 30sLint / static analysis
    medpush
  6. 1minRuntime schema validation
    highpush
  7. 1minPair / mob programming
    medpush
  8. 1minSecrets detection
    highpush
  9. 2minUnit tests
    highpush
  10. 5minAI code review
    medpush
  11. 10minCoverage / mutation testing
    medpush
  12. 10minIntegration tests
    highpush
  13. 10minVisual regression testing
    medpush
  14. 30minE2E tests
    medpush
  15. 30minSAST / code security scan
    medpush
  16. 1hrError monitoring
    highpush
  17. 12hrCode review
    medpush
  18. 1wkBug reports
    medpush

Is production healthy right now?

These loops have to be always on, and typically have to be pushed to you so that if something isn't behaving within the right bounds, you can intervene without constantly looking at dashboards. Take care that the checks are push. If they are pull, customers will be telling you about outages.

  1. 5minSynthetic monitoring / uptime checks
    medpush
  2. 15minAlerts / on-call pages
    highpush
  3. 1hrRuntime security detection
    medpush
  4. 1hrAPM / distributed tracing
    highpull
  5. 1hrMemory / resource utilization
    medpush
  6. 1hrError monitoring
    highpush

Can we keep shipping safely?

If we neglect parts of the development process, things can start breaking more. It might be we have a tendency to miss something before a deploy, or the issues may be subtler. If we are getting signal back from these loops, it's a sign there's a gap earlier in the SDLC.

  1. 30minCanary / progressive rollout
    highpush
  2. 1dChaos engineering / game days
    highpull
  3. 1wkSLO / error budget burn
    highpush
  4. 1moChange failure rate
    medpull

Could an adversary get in, and would we know?

You can never be certain that a mistake won't be introduced, but you can do all you can to make sure you are staying on secure dependencies and that you don't slip up (or that you find out about your slip-ups before an adversary does).

  1. 1minSecrets detection
    highpush
  2. 30minSAST / code security scan
    medpush
  3. 1hrRuntime security detection
    medpush
  4. 1dThreat modeling
    medpull
  5. 1dSCA / dependency vulnerabilities
    medpush
  6. 1dCloud misconfiguration / CSPM
    highpush
  7. 1wkDAST / automated web scanning
    medpush
  8. 2wkBug bounty / VDP
    highpush
  9. 2wkAI red-teaming / adversarial evals
    highpull
  10. 1moOpen vulnerability age / patch SLA
    highpull
  11. 3moPen test / red team
    highpull

What broke, and what do we learn from it?

We want to make sure that we learn from anything that goes wrong, and decrease the time it takes to recover. Ideally we look forward at the shortcomings of our system to work out where we could be exposed.

  1. 1dPre-mortems
    lowpull
  2. 1wkIncident post-mortems
    highpull
  3. 1moMean time to restore
    medpull

Is it fast enough for real users?

End-user performance is exceptionally important for take-up and a pleasurable user experience. Without keeping a close eye on this, it is very likely we'll introduce performance regressions as the system becomes more capable.

  1. 3minBundle size / asset weight
    highpush
  2. 10minLighthouse / synthetic performance
    highpush
  3. 30minBenchmark tests
    medpush
  4. 1hrModel latency / time-to-first-token
    highpull
  5. 1wkLoad / stress testing
    medpull
  6. 1moCore Web Vitals (field / RUM)
    highpull
  7. 1moCapacity / scaling signals
    medpush

Is our delivery machine healthy?

Often people optimise for resource efficiency (higher utilisation). This works against flow and responsiveness; the overall delivery machine is much more efficient when we drive down batch size and the number of things we work on at once, and permit some amount of slack in the system so items can move through the process effectively.

  1. 1dDaily standup
    medpush
  2. 1dWIP
    highpull
  3. 1dBuild & CI duration
    medpull
  4. 2dApp store review gate
    medpush
  5. 3dWork item age
    medpull
  6. 2wkTask cycle time
    medpull
  7. 2wkThroughput
    medpull
  8. 1moPR cycle time
    medpull
  9. 1moDeployment frequency
    medpull
  10. 1moLead time for changes
    medpull
  11. 1moSprint predictability (say:do ratio)
    lowpull
  12. 1moFlow efficiency
    medpull

Where should we invest in the codebase?

If we're often reworking the same piece of code, it's a sign that one area of code has multiple reasons for change, and there may be an opportunity to tease apart those multiple concerns into separate components. It is gratifying on the occasions when this makes the codebase more capable than it was.

  1. 5minUnused code / dead exports
    medpull
  2. 10minCode complexity & coupling metrics
    medpush
  3. 1wkRework rate / code churn
    medpull
  4. 1moHotspot analysis (churn × complexity)
    highpull
  5. 1moDependency freshness (libyear)
    highpull
  6. 1yrTech debt accumulation
    lowpull

Can we even build this?

Sometimes we simply can't tell off the top of our head whether something will work, and we need to get a bit further along. This gives us sharper visibility.

  1. 1dFeasibility spike / throwaway prototype
    highpull
  2. 1wkTechnology / model evaluation
    medpull
  3. 1wkDesign / RFC review
    medpull

Can downstream consumers trust the data?

It may be that our abstraction mistakenly conflates data modelling and data transport, or we have an error with codecs/marshalling. If we're interacting with important systems to deliver our service to customers, we can improve reliability by automatically testing the requirements of those systems.

  1. 15minData quality tests
    highpush
  2. 1hrPipeline freshness / data SLAs
    highpush
  3. 1dData anomaly detection
    medpush

Is the AI feature actually good?

The non-deterministic nature of LLM-based features introduces new failure modes. If a model is changed or quantised, this can wreck useful functionality. Where several parts make up a prompt, changing one can impact the whole result. If input or prompt data can be influenced by the customer, we can get completely unanticipated results due to the way the parts of the prompt interact.

  1. 30minEval suite / golden-set regression
    medpush
  2. 1hrToken / cost per request
    highpull
  3. 1hrLLM-as-judge scoring
    medpull
  4. 1dThumbs up/down on AI output
    lowpull
  5. 1wkModel & output drift
    medpull
  6. 2wkAI red-teaming / adversarial evals
    highpull

Value

Is it worth making?

Should we build this at all?

It is very easy to fall in love with a solution. I've lost count of the number of times I've done it. But the best thing you can do is treat your solution as a series of theories or bets that you have to validate. The moment you interact with real people (or find out that you can't find them), you'll get feedback on those theories.

  1. 1wkFake door / prototype tests
    medpull
  2. 1wkWeekly discovery interviews
    highpull
  3. 1wkFeature request board / upvotes
    medpull
  4. 2wkKano survey
    medpull
  5. 2wkOpportunity scoring (importance vs satisfaction)
    medpull
  6. 2wkWillingness-to-pay research
    medpull

Can people actually use it?

Our own mental models of how software should be used often don't correspond to the mental models that our users have. When we move to implementation, our usability then has 'gotchas' that aren't as visible to us. This is another area where it's easy to fall in love with a solution, perhaps because it is visually appealing compared to what we had before, or it neatly mirrors what is happening technically. But that's different to meeting as many users as possible where they are.

  1. 10minAccessibility checks
    medpush
  2. 1dDogfooding
    medpush
  3. 1dFirst-click / findability
    highpull
  4. 1dClicks / interaction cost
    medpull
  5. 1dTask ease (SEQ / SUS)
    medpull
  6. 1dSession recordings / rage clicks
    medpull
  7. 1wkBeta / early-access program
    highpull
  8. 1wkUsability tests
    highpull
  9. 1wkTask success rate
    highpull
  10. 1wkTime on task
    medpull
  11. 1wkFunnel / drop-off analysis
    highpull

Did the thing we shipped actually work?

After we've delivered something, if people don't use it, was this because it wasn't communicated well? Or is it that we were incorrect about the appeal of what we delivered? If people use it, do they get the outcomes they are intended to get?

  1. 1wkFeature adoption / activation
    medpull
  2. 1wkFunnel / drop-off analysis
    highpull
  3. 2wkSprint review / stakeholder demo
    medpush
  4. 2wkA/B experiments
    highpull

Do we have product-market fit, and are we keeping it?

Sometimes great initial conversations or marketing and sales motion can disguise the appeal or use of a system, and as we fail to make customers happy or to retain them, we see that what we're doing doesn't have a strong trajectory in the right direction.

  1. 1dAI synthesis of qualitative streams
    medpull
  2. 1wkPublic reviews (G2 / app stores / Reddit)
    medpull
  3. 1wkChurn interviews / exit surveys
    highpull
  4. 2wkSean Ellis PMF survey
    medpull
  5. 1moNPS / CSAT
    lowpull
  6. 1moStickiness (DAU/MAU)
    medpull
  7. 1moTrial / freemium conversion
    highpull
  8. 3moRetention curves
    highpull
  9. 6moExpansion / referral revenue
    highpull

Are we priced right?

Our price may not be what we need it to be. If we can find out why, we can hopefully see whether that's to do with our positioning, the expectations of our market, what value our product delivers, or how we communicate that value.

  1. 1wkChurn interviews / exit surveys
    highpull
  2. 2wkWillingness-to-pay research
    medpull
  3. 1moTrial / freemium conversion
    highpull
  4. 1moWin/loss analysis
    highpull
  5. 3moPricing realization / discount rate
    highpull
  6. 1yrPricing power
    medpull

Is support scaling with the product?

Ideally we have a process that surfaces and prioritises for the product team the items that create the most burden on support. We also want to make sure that our support function is delivering what it needs to, to help customers achieve what they need.

  1. 4hrTime to first response
    medpull
  2. 1dResolution time
    medpull
  3. 1wkTicket CSAT
    lowpull
  4. 1wkSupport tickets
    medpush
  5. 2wkFirst contact resolution
    medpull
  6. 1moContact rate / tickets per customer
    medpull

Reach

Can we get it to people?

Which channel deserves the next dollar?

The fast numbers aren't very useful without the slow numbers. It's very tempting to try to drive down a cost, but it might be that customers that are cheaper to acquire also give you less value, so care has to be taken to not optimise for one thing.

  1. 1dCPM / auction pressure
    medpull
  2. 1dLast-click attribution
    lowpull
  3. 1wkAd-spend liquidation
    highpull
  4. 1wkSelf-reported attribution
    medpull
  5. 1wkMulti-touch / data-driven attribution
    medpull
  6. 1wkInfluencer / sponsorship performance
    lowpull
  7. 1wkEvents / webinar performance
    medpull
  8. 1moROAS / channel CAC
    medpull
  9. 1moIncrementality / lift tests
    highpull
  10. 3moPartnership pipeline
    medpull
  11. 3moBlended CAC
    medpull
  12. 6moMarketing mix modeling
    lowpull

Does this message land with this audience?

These figures are essentially feedback on how effective your creative is, whether through organic social or ads. Did you stop the scroll? How effective was your hook? Do your open loops keep viewers? Are you delivering the payoff? Does this attract regular followers/viewers?

  1. 4hrOpen rate
    lowpull
  2. 4hrClick-through rate
    medpull
  3. 6hrViews / impressions
    lowpull
  4. 6hrComments / audience conversation
    medpull
  5. 6hrComments & DMs
    medpull
  6. 12hrCompletion / loop rate
    highpull
  7. 1dHook rate / swipe-aways
    medpull
  8. 1dAudience retention / watch time
    highpull
  9. 1dFollows-per-view
    medpull
  10. 1dPaid ads CTR / CPA
    medpull
  11. 1dCost per click (CPC)
    medpull
  12. 1dEngagement rate
    medpull
  13. 3dLanding conversion rate (CVR)
    medpull
  14. 3dContent engagement
    lowpull
  15. 3dCold outreach reply rates
    medpush
  16. 1wkQuality / relevance score
    medpull
  17. 2wkFrequency / creative fatigue
    medpull

Can we still reach the inbox?

This has two areas of relevance, one for day-to-day operations: if we mess up our reputation with our transactional API (such as Amazon SES), it can cause a lot of pain. There's also cold bulk email against leads, the sharp end, where it is important to use multiple domains and protect your main domain.

  1. 5minSpam / reject rate
    medpull
  2. 10minBounce rate
    highpull
  3. 15minDelivered / inbox placement
    highpull
  4. 1dUnsubscribe / list churn
    highpull
  5. 1wkSender reputation
    medpull

Are we compounding an audience or renting one?

Advertising with a simple X in, Y out relationship can lead to incredible growth for as long as we can support that cash flow, but we don't want to neglect channels where we can multiply our ongoing results each time we invest in them.

  1. 1dShares & saves
    highpull
  2. 1dFollows-per-view
    medpull
  3. 1dCommunity activity & sentiment
    medpull
  4. 2wkPR / press coverage
    lowpull
  5. 1moFollower / subscriber growth
    medpull
  6. 1moReferral / K-factor
    medpull
  7. 1yrBrand awareness / recall
    lowpull

Can buyers find us when they go looking?

SEO as a channel has a huge advantage: you are carving out a stream of leads, if you are capable of getting ranked (which depends on how competitive the space is, and how people search for it). But it has the huge disadvantage that it takes a long time to get actual leads. If you are pre-PMF, this may mean that you are trying to get ranked for the wrong thing.

  1. 1dKeyword research (difficulty vs. potential)
    medpull
  2. 1wkTechnical SEO / indexation
    highpull
  3. 1moKeyword rankings / SERP position
    medpull
  4. 1moOrganic traffic & CTR
    highpull
  5. 1moAI search visibility (AEO / GEO)
    lowpull
  6. 1moApp store ratings & ASO
    medpull
  7. 3moDomain authority / backlink profile
    medpull
  8. 6moContent decay / refresh signal
    medpull

Is demand converting into revenue?

We want to distinguish whether a stage in the pipeline is performing badly for mechanical reasons that we can fix, or whether we've got a more fundamental issue of positioning or lack of perceived value.

  1. 1moMeetings booked / pipeline
    medpull
  2. 1moWin/loss analysis
    highpull
  3. 3moPipeline stage conversion
    medpull
  4. 3moSales cycle length
    medpull

Team

Can we sustain the people doing the thing?

Is the hiring machine working?

Depending on when and how many people we need, some sources may not sustain the candidate flow that we need. We also want to instrument our own process. Some questions to ask ourselves: Are the things we are asking for actually giving the signal we need? Can we reduce the number of stages? The slower one moves, the more likely we are to lose awesome candidates.

  1. 3dSourcing response rate
    medpull
  2. 1wkInterview signal
    lowpull
  3. 1wkOnboarding ramp / time-to-first-commit
    medpull
  4. 1moOffer acceptance rate
    medpull
  5. 1moTime-to-hire
    medpull
  6. 6moQuality of hire / ramp time
    medpull
  7. 6moCompensation benchmarking
    medpull

Is the team healthy and sustainable?

This area is very, very human. Our own biases can stop us seeing clearly how other people perceive situations. In the end we want to have a workplace, team and organisation that are invigorating and enjoyable.

  1. 1wk1:1 sentiment
    medpull
  2. 1wkPulse surveys
    medpull
  3. 1wkMeeting load / focus time
    medpull
  4. 1wkOn-call load / pages per person
    highpull
  5. 2wkTeam retrospectives
    medpull
  6. 1moDeveloper experience survey
    medpull
  7. 3moeNPS / engagement surveys
    lowpull
  8. 3moBus factor / knowledge concentration
    medpull
  9. 1yrAttrition / regretted departures
    highpush

Are people growing here?

If anything here is a surprise for the person receiving feedback, then it is probable we haven't done a good enough job of communicating in the intervening period.

  1. 3moPerformance reviews / 360 feedback
    medpull
  2. 6moGrowth / promotion readiness
    medpull

Are we working on the right things?

If the company already has a specific objective, are we aligned with it? If we are trying to find where we can add value, are we considering our options?

  1. 2wkOpportunity scoring (importance vs satisfaction)
    medpull
  2. 3moOKR check-ins / goal scoring
    medpull

Risk & Moat

Do the economics allow success to continue?

Do the unit economics actually work?

The basics: are we spending less on our AI than we're getting back, and are we spending less to get a customer than we're getting back?

  1. 1wkAI / inference spend
    highpush
  2. 3moGross margin
    highpull
  3. 3moBlended CAC
    medpull
  4. 3moCAC payback period
    medpull
  5. 3moPricing realization / discount rate
    highpull
  6. 1yrLTV
    lowpull
  7. 1yrLTV:CAC ratio
    lowpull

How long can we keep playing?

Money in versus money out, and therefore how many months before we have to raise or cut.

  1. 1dCloud spend / burn alerts
    highpush
  2. 1wk13-week cash flow forecast
    highpull
  3. 1moBudget vs actuals (monthly close)
    highpull
  4. 1moCash collection / DSO
    highpull
  5. 1moInvestor / fundraising feedback
    lowpull
  6. 3moRunway / burn multiple
    highpull

Is revenue healthy, not just growing?

Churn/retention can kill things even if top-line growth appears impressive, so make sure you are looking at a balanced set of signals.

  1. 3moMRR / ARR growth
    highpull
  2. 3moRevenue churn / GRR
    highpull
  3. 3moRevenue concentration
    highpull
  4. 3moRevenue forecast accuracy
    medpull
  5. 6moExpansion / referral revenue
    highpull

What could kill us from outside?

Unfortunately nothing consistent will tell us about the varied activities of platforms, vendors, regulators, fraudsters and competitors.

  1. 1dFraud / abuse / chargeback rate
    medpush
  2. 1moPlatform / API dependency changes
    lowpull
  3. 1moCompetitor launches
    lowpull
  4. 3moVendor / provider concentration
    lowpull
  5. 3moCompliance audits (SOC2 / ISO)
    highpull
  6. 6moRegulatory & legal signals
    medpull
  7. 1yrChurn by tenure / lock-in strength
    medpull
  8. 1yrPricing power
    medpull

Dataset v0.23.1 · 204 loops · CC BY 4.0 · cite as: Konrad Bloor, Feedback Loop Field Guide, konradbloor.com/loops