<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Bruno Scheufler]]></title><description><![CDATA[building startups and distributed systems]]></description><link>https://www.brunoscheufler.com</link><image><url>https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png</url><title>Bruno Scheufler</title><link>https://www.brunoscheufler.com</link></image><generator>Substack</generator><lastBuildDate>Mon, 21 Sep 2026 20:32:28 GMT</lastBuildDate><atom:link href="https://www.brunoscheufler.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Bruno Scheufler]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[brunoscheufler@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[brunoscheufler@substack.com]]></itunes:email><itunes:name><![CDATA[Bruno Scheufler]]></itunes:name></itunes:owner><itunes:author><![CDATA[Bruno Scheufler]]></itunes:author><googleplay:owner><![CDATA[brunoscheufler@substack.com]]></googleplay:owner><googleplay:email><![CDATA[brunoscheufler@substack.com]]></googleplay:email><googleplay:author><![CDATA[Bruno Scheufler]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[My Next Chapter]]></title><description><![CDATA[Today, I'm moving to San Francisco.]]></description><link>https://www.brunoscheufler.com/p/2026-07-05-my-next-chapter</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2026-07-05-my-next-chapter</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 05 Jul 2026 19:00:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4456949b-c958-4901-ae28-d207365eb015_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!h5qf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!h5qf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!h5qf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!h5qf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!h5qf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!h5qf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;My Next Chapter&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="My Next Chapter" title="My Next Chapter" srcset="https://substackcdn.com/image/fetch/$s_!h5qf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!h5qf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!h5qf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!h5qf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1436ccfd-383e-4d34-960a-3aec2a80ea15_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><p>Today, I'm moving to San Francisco. Moving to the United States to work on a startup has been one of my biggest dreams ever, and it became reality when my O-1 was approved shortly after my 25th birthday two months ago. I've been working on this for more than a decade, so today, I'd like to share my whole story up to this point.</p><p>I was born in Freiburg, a city in southern Germany. My parents are doctors and have always worked hard. Because of their jobs, we moved frequently, I grew up in Zurich, Switzerland and Innsbruck, Austria. I've had an interest in building since I was a kid. At twelve years old, I started extending my favorite video game to play with friends. Over time, I moved on to mobile and web application development, using any free minute I could get to learn more and work hard to get better. Little did I know that my fun hobby would turn into a career.</p><p>I've always been drawn to startups. Autonomy and ownership are two of the most important things I value. I went to high school close to Frankfurt and got my first summer internship as a software engineer in 2017, when I was 15 years old. I loved the experience so much that I swore to myself to learn harder than ever before to get a real job. A year later, I <a href="https://brunoscheufler.com/2019-05-12-whats-next/">landed my first job</a> as a software engineer at Hygraph. I worked harder than ever before, and joined the company full-time after graduating high school in 2019, watching the company grow from five people to over 70 in the span of a couple years.</p><p>During Covid, I moved to Munich where I've lived for the past five years. I studied at TUM and LMU, the two most prestigious universities in Germany, graduating from TUM in 2023. I tried building a startup with my best friend and even though it didn't work out at the time, I learned valuable lessons about the business environment and challenges software startups are still facing in Germany.</p><p>I took everything I learned up to that point and <a href="https://brunoscheufler.com/2024-04-23-joining-inngest/">joined Inngest</a> as a distributed systems engineer in early 2024. Through Inngest, I was able to visit San Francisco three times in the past two years and I fell in love with the city. Tech runs deep in San Francisco (you really can't escape it) and the energy and level of optimism in everyone I've met so far are unlike anywhere else. People work incredibly hard and builders are taken seriously, with lots of trust and support. In December 2024, I knew that I wanted to move to the US and I started the visa process that would take more than 18 months.</p><p>If I had told my fifteen-year-old self that within ten years, I would be living and working for a startup in Silicon Valley, I simply wouldn't have believed it. This is the power of pursuing a dream, joined with the opportunities presented by the US. I learned the value of working hard and never giving up, no matter how many roadblocks you find in your way. I also learned the value of focusing on what matters most and understanding opportunity cost. Time and attention are the scarcest resources in life.</p><p>Moving to the US is just the beginning. I want to work hard, learn from the best, grow, and help shape the future with products that have a real impact on people's lives. I deeply believe that now is the best time to build, in all of human history. The tools and resources we have at our disposal have never been more powerful. The question is no longer how to build, it's what to build, and why. The challenge is to assemble a world-class team and build the best possible product out there to solve real problems.</p><p>I'm incredibly excited for the next chapter and I'll keep sharing more on my journey.</p><p>I want to thank everyone without whom I wouldn't be where I am today.</p><p>First of all, I want to thank my amazing girlfriend for being supportive of this plan. I truly couldn't have done this without you. I want to thank my good friends Tim, Mohamad, Kai, Nico, Hendrik, Timo, Benjamin, Noelia, Paul, Lorenz, Miguel, and Philipp. You've always encouraged me to challenge myself and dream bigger.</p><p>From Inngest, I want to thank the entire team, with special thanks to Tony, Darwin, Albert, Jacob, Jakob, Lakshmi, Riadh, and Muzammil. Thank you for your trust and belief in me.</p><p>From Hygraph, I want to thank Michael, Daniel, Jonas, Fabi, Dino, Pablo, Jean, Julian, Frederik, and Larisa. Thank you for teaching me more in a couple years than I ever believed was possible.</p><p>From TUM, I want to thank Prof. Pramod Bhatotia for your support during my thesis and the research projects you invited me to help on.</p><p>I also want to thank Jennifer Li, Martin Casado, and Yoko Li from a16z, Ivan from Daytona, Andrea from Mendral, Jake from Railway, and Mathias from Netlify for supporting me with the visa.</p><p><em>Every second counts</em></p>]]></content:encoded></item><item><title><![CDATA[Looking back on 2025]]></title><description><![CDATA[As another eventful year is coming to an end, I'm continuing my annual tradition of reflecting on 2025.]]></description><link>https://www.brunoscheufler.com/p/2025-12-31-looking-back-on-2025</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2025-12-31-looking-back-on-2025</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Wed, 31 Dec 2025 13:41:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As another eventful year is coming to an end, I'm continuing my <a href="https://brunoscheufler.com/2024-12-30-looking-back-on-2024/">annual tradition</a> of reflecting on 2025. Last year was marked by significant change: I <a href="https://brunoscheufler.com/2024-04-23-joining-inngest/">joined Inngest</a> in April 2024, which has come to be one of the best decisions I've ever made.</p><p>This year, I got to work on numerous challenging and exciting projects, including rewriting nearly all of the internal systems ingesting billions of events and powering hundreds of millions of workflow runs every day. We've raised a <a href="https://www.inngest.com/blog/announcing-inngest-series-a?ref=brunoscheufler.com">$21m Series A</a> earlier this year and are facing the best (tons of happy customers with amazing use cases) and worst (scaling challenges) effects of product market fit.</p><p>This August, I traveled to London to speak at <a href="https://www.gophercon.co.uk/?ref=brunoscheufler.com">GopherCon UK</a> about the lessons I learned migrating mission-critical systems with zero downtime. This was my first talk in years, so I was reasonably nervous leading up to it. Judging from the feedback I got by attendees walking up to me after the session, I did better than I anticipated and I'm looking forward to return for another round next year!</p><p>In October, we held our annual team offsite in Boston. It was inspiring to see how much the team has grown this year (we more than doubled in size), and I can't wait to see everyone in person again soon. I also got to travel to San Francisco on numerous occasions and I always return more energized than I left.</p><p>As we're entering the last days of 2025, I'm recharging and looking ahead on to 2026. I'm confident that this will be the best year yet.</p>]]></content:encoded></item><item><title><![CDATA[Bidirectional Markdown syncing for Ghost]]></title><description><![CDATA[A couple days ago, I pulled the trigger and moved this blog from Next.js to Ghost.]]></description><link>https://www.brunoscheufler.com/p/2025-12-22-bidirectional-markdown-syncing-for-ghost</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2025-12-22-bidirectional-markdown-syncing-for-ghost</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Mon, 22 Dec 2025 00:42:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A couple days ago, I pulled the trigger and moved this blog from Next.js to Ghost. You can read more about the process behind this decision <a href="https://brunoscheufler.com/2025-12-22-on-picking-software/">in my other piece</a>. In this post, I want to focus on <em>how</em> the new stack works, and why I think this should stand the test of time. Here goes.</p><p>My primary focus for the architecture of this blog always has been the content backbone. Post content has been stored as Markdown files in a Git repository for the past 8 years and I do not expect this to change in the next 8 years either. Images and media are also stored in the repository using Git LFS, and hosted on my static subdomain with proper Cache-Control headers.</p><p>Previously, I spent a lot of effort on designing a polished landing page, including animations "inspired" by the best (are you really a web developer until you've attempted to copy Stripe?). For this iteration, my focus was on the content. As long as I found a way to get the Markdown files properly displayed, I'd be happy.</p><p>My good friend and frequent collaborator <a href="https://timweiss.net/?ref=brunoscheufler.com">Tim</a> switched his blog to <a href="https://ghost.org/?ref=brunoscheufler.com">Ghost</a>, an open content and newsletter platform you can self-host. I trust his taste in software and share his idea about ownership, so I didn't think long before making this decision.</p><p>Self-hosting Ghost isn't a miracle, and their <a href="https://docs.ghost.org/hosting?ref=brunoscheufler.com">documentation</a> contains everything you need. Where it gets a little more interesting is my way of syncing content: Most people probably just opt to author their posts in Ghost's fantastic editor (as did I, for the post you're reading at this very moment), yet I wanted the guarantee of having all my content on disk, neatly versioned and all.</p><h2>Syncing from Markdown to Ghost</h2><p>To get my 230+ existing posts onto Ghost in the first place, I had to write a script to convert the Markdown source to HTML and upload everything via the built-in <a href="https://docs.ghost.org/admin-api?ref=brunoscheufler.com">Admin API</a>. This was easier than I thought, and with the help of Claude Code, I got to the following solution:</p><pre><code>#!/usr/bin/env python3
"""
Ghost CMS Sync Script
Syncs Markdown blog posts from content/posts/ to Ghost CMS.
"""

import os
import sys
import time
import glob as file_glob
from datetime import datetime, timezone
from pathlib import Path
from typing import Dict, List, Optional, Tuple

import jwt
import requests
import frontmatter
import markdown2
from dotenv import load_dotenv


# Debug mode - set to True to see HTML output
DEBUG = os.getenv('DEBUG', 'false').lower() == 'true'


# Configuration
GHOST_API_URL = "https://YOUR_URL_HERE"
CONTENT_DIR = Path(__file__).parent.parent / "content" / "posts"
ENV_FILE = Path(__file__).parent.parent / ".env"

# External post types that should not be synced to Ghost
EXTERNAL_POST_TYPES = ['EXTERNAL']


class GhostAPIClient:
    """Client for interacting with Ghost Admin API."""

    def __init__(self, api_url: str, admin_api_key: str):
        self.api_url = api_url.rstrip('/')
        self.admin_api_key = admin_api_key

        # Parse the admin API key
        try:
            self.key_id, self.key_secret = admin_api_key.split(':')
        except ValueError:
            raise ValueError("Invalid GHOST_ADMIN_API_KEY format. Expected 'ID:SECRET'")

    def _generate_jwt_token(self) -&gt; str:
        """Generate a JWT token for Ghost Admin API authentication."""
        iat = int(datetime.now(timezone.utc).timestamp())

        header = {
            'alg': 'HS256',
            'typ': 'JWT',
            'kid': self.key_id
        }

        payload = {
            'iat': iat,
            'exp': iat + 300,  # Token expires in 5 minutes
            'aud': '/admin/'
        }

        token = jwt.encode(payload, bytes.fromhex(self.key_secret), algorithm='HS256', headers=header)
        return token

    def _make_request(self, method: str, endpoint: str, data: Optional[Dict] = None) -&gt; requests.Response:
        """Make an authenticated request to Ghost API."""
        token = self._generate_jwt_token()
        url = f"{self.api_url}{endpoint}"

        headers = {
            'Authorization': f'Ghost {token}',
            'Content-Type': 'application/json'
        }

        response = requests.request(method, url, json=data, headers=headers)
        return response

    def get_post_by_slug(self, slug: str) -&gt; Optional[Dict]:
        """Get a post by its slug."""
        try:
            response = self._make_request('GET', f'/ghost/api/admin/posts/slug/{slug}/')
            if response.status_code == 200:
                return response.json()['posts'][0]
            elif response.status_code == 404:
                return None
            else:
                print(f"  Warning: Unexpected status {response.status_code} when fetching slug {slug}")
                return None
        except Exception as e:
            print(f"  Error fetching post by slug {slug}: {e}")
            return None

    def get_posts_by_title(self, title: str) -&gt; List[Dict]:
        """Get all posts with a specific title."""
        try:
            # Use filter parameter to search by title
            response = self._make_request('GET', f'/ghost/api/admin/posts/?filter=title:\'{title}\'&amp;limit=all')
            if response.status_code == 200:
                return response.json()['posts']
            else:
                print(f"  Warning: Unexpected status {response.status_code} when searching for title")
                return []
        except Exception as e:
            print(f"  Error searching posts by title: {e}")
            return []

    def delete_post(self, post_id: str) -&gt; Tuple[bool, Optional[str]]:
        """Delete a post from Ghost."""
        try:
            response = self._make_request('DELETE', f'/ghost/api/admin/posts/{post_id}/')
            if response.status_code == 204:
                return True, None
            else:
                error_msg = response.json().get('errors', [{}])[0].get('message', 'Unknown error')
                return False, f"Status {response.status_code}: {error_msg}"
        except Exception as e:
            return False, str(e)

    def create_post(self, post_data: Dict) -&gt; Tuple[bool, Optional[str]]:
        """Create a new post in Ghost."""
        try:
            response = self._make_request('POST', '/ghost/api/admin/posts/?source=html', {'posts': [post_data]})
            if response.status_code in [200, 201]:
                return True, None
            else:
                error_msg = response.json().get('errors', [{}])[0].get('message', 'Unknown error')
                return False, f"Status {response.status_code}: {error_msg}"
        except Exception as e:
            return False, str(e)

    def update_post(self, post_id: str, post_data: Dict) -&gt; Tuple[bool, Optional[str]]:
        """Update an existing post in Ghost."""
        try:
            response = self._make_request('PUT', f'/ghost/api/admin/posts/{post_id}/?source=html', {'posts': [post_data]})
            if response.status_code == 200:
                return True, None
            else:
                error_msg = response.json().get('errors', [{}])[0].get('message', 'Unknown error')
                return False, f"Status {response.status_code}: {error_msg}"
        except Exception as e:
            return False, str(e)


def parse_markdown_file(file_path: Path) -&gt; Optional[Tuple[Dict, str]]:
    """Parse a markdown file and return frontmatter and content."""
    try:
        with open(file_path, 'r', encoding='utf-8') as f:
            post = frontmatter.load(f)

        # Validate required fields
        required_fields = ['title', 'path', 'date', 'published']
        missing_fields = [field for field in required_fields if field not in post.metadata]

        if missing_fields:
            print(f"  &#9888; Skipping {file_path.name}: Missing fields {missing_fields}")
            return None

        return post.metadata, post.content
    except Exception as e:
        print(f"  &#9888; Error parsing {file_path.name}: {e}")
        return None


def convert_markdown_to_html(markdown_content: str) -&gt; str:
    """Convert Markdown content to HTML, preserving code block language classes."""
    import re
    import html as html_lib

    # Extract fenced code blocks with language info
    code_blocks = []

    def extract_fenced_block(match):
        language = match.group(1) or ''
        code = match.group(2)
        # Use HTML comment as placeholder (won't be processed by markdown2)
        placeholder = f'&lt;!--CODE_BLOCK_{len(code_blocks)}--&gt;'
        code_blocks.append((language, code))
        return placeholder

    # Extract fenced code blocks (with or without language)
    fenced_pattern = r'```(\w+)?\n(.*?)```'
    temp_md = re.sub(fenced_pattern, extract_fenced_block, markdown_content, flags=re.DOTALL)

    # Convert markdown to HTML (without fenced-code-blocks extra since we handle it manually)
    # Note: 'code-friendly' removed to enable underscore-based emphasis (_italic_)
    extras = ['tables', 'break-on-newline']
    html = markdown2.markdown(temp_md, extras=extras)

    # Replace placeholders with properly formatted code blocks
    for i, (language, code) in enumerate(code_blocks):
        placeholder = f'&lt;!--CODE_BLOCK_{i}--&gt;'
        code_escaped = html_lib.escape(code.rstrip())

        if language:
            replacement = f'&lt;pre&gt;&lt;code class="language-{language}"&gt;{code_escaped}&lt;/code&gt;&lt;/pre&gt;'
        else:
            replacement = f'&lt;pre&gt;&lt;code&gt;{code_escaped}&lt;/code&gt;&lt;/pre&gt;'

        html = html.replace(placeholder, replacement)

    return html


def extract_slug_from_path(path: str) -&gt; str:
    """Extract slug from the path field."""
    # Handle full URLs (external posts)
    if path.startswith('http://') or path.startswith('https://'):
        # Extract the last segment from the URL path
        # e.g., 'https://www.inngest.com/blog/sharding-at-inngest' -&gt; 'sharding-at-inngest'
        from urllib.parse import urlparse
        parsed = urlparse(path)
        slug = parsed.path.rstrip('/').split('/')[-1]
        return slug if slug else path

    # Handle local paths like '/blog/2024-12-30-post-title'
    # The path is like '/blog/2024-12-30-looking-back-on-2024'
    # We want to extract just '2024-12-30-looking-back-on-2024'
    # Ghost will add the /blog/ prefix via its routing configuration
    if path.startswith('/blog/'):
        return path[6:]  # Remove '/blog/' prefix

    # Fallback for other formats
    return path.lstrip('/')


def map_to_ghost_post(frontmatter: Dict, html_content: str) -&gt; Dict:
    """Map frontmatter and content to Ghost post format."""
    slug = extract_slug_from_path(frontmatter['path'])

    # Map published boolean to status
    status = 'published' if frontmatter.get('published', True) else 'draft'

    # Get topics as tags (or empty list if not present)
    tags = frontmatter.get('topics', [])

    # Ensure tags is a list
    if not isinstance(tags, list):
        tags = [tags]

    # Generate keywords meta tag for code injection
    keywords = frontmatter.get('keywords', [])
    keywords_tag = None
    if keywords and isinstance(keywords, list) and len(keywords) &gt; 0:
        # Filter out None values and ensure all keywords are strings
        keywords = [str(k) for k in keywords if k is not None]
        if keywords:  # Only create meta tag if there are valid keywords
            # Escape special characters for HTML attributes
            import html
            keywords_str = ', '.join(keywords)
            keywords_escaped = html.escape(keywords_str, quote=True)
            keywords_tag = f'&lt;meta name="keywords" content="{keywords_escaped}" /&gt;'

    # Generate og:image meta tag for code injection
    ogimage_tag = None
    if 'ogImageUrl' in frontmatter and frontmatter['ogImageUrl']:
        import html
        ogimage_url = html.escape(str(frontmatter['ogImageUrl']), quote=True)
        ogimage_tag = f'&lt;meta property="og:image" content="{ogimage_url}" /&gt;'

    ghost_post = {
        'title': frontmatter['title'],
        'slug': slug,
        'html': html_content,
        'status': status,
        'published_at': frontmatter['date'],
        'tags': tags
    }

    # Add feature image if present
    if 'headerImageUrl' in frontmatter:
        ghost_post['feature_image'] = frontmatter['headerImageUrl']
        ghost_post['feature_image_alt'] = frontmatter['title']

    # Add code injection if keywords or og:image exist
    codeinjection_tags = []
    if keywords_tag:
        codeinjection_tags.append(keywords_tag)
    if ogimage_tag:
        codeinjection_tags.append(ogimage_tag)

    if codeinjection_tags:
        ghost_post['codeinjection_head'] = '\n'.join(codeinjection_tags)

    return ghost_post


def sync_post_to_ghost(client: GhostAPIClient, file_path: Path, dry_run: bool = False, cleanup_duplicates: bool = False) -&gt; str:
    """Sync a single post to Ghost. Returns status: 'created', 'updated', 'skipped', or 'error'."""
    # Parse the markdown file
    result = parse_markdown_file(file_path)
    if result is None:
        return 'error'

    frontmatter_data, markdown_content = result

    # Skip external post types
    post_type = frontmatter_data.get('type', '')
    if post_type in EXTERNAL_POST_TYPES:
        print(f"  &#8856; Skipping external post type: {post_type}")
        return 'skipped'

    # Convert markdown to HTML
    html_content = convert_markdown_to_html(markdown_content)

    # Map to Ghost post format
    ghost_post = map_to_ghost_post(frontmatter_data, html_content)
    slug = ghost_post['slug']
    title = ghost_post['title']

    # Debug output
    if DEBUG:
        print(f"\n  Slug: {slug}")
        print(f"  Title: {ghost_post['title']}")
        print(f"  Status: {ghost_post['status']}")
        print(f"  Tags: {ghost_post['tags']}")
        print(f"  HTML preview (first 500 chars):\n{html_content[:500]}\n")
        if '&lt;img' in html_content:
            print(f"  &#10003; Images found in HTML")
            # Extract and show image tags
            import re
            img_tags = re.findall(r'&lt;img[^&gt;]+&gt;', html_content)
            for img in img_tags[:3]:  # Show first 3 images
                print(f"    {img}")
        else:
            print(f"  &#10007; No images found in HTML")

    if dry_run:
        print(f"  [DRY RUN] Would sync: {slug}")
        # Check for duplicates in dry-run mode too
        posts_with_title = client.get_posts_by_title(title)
        if len(posts_with_title) &gt; 1:
            print(f"  &#9888; Found {len(posts_with_title)} posts with title '{title}':")
            for p in posts_with_title:
                print(f"    - {p['slug']} (ID: {p['id']}, updated: {p['updated_at'][:10]})")
        return 'skipped'

    # Check if post exists by slug
    existing_post = client.get_post_by_slug(slug)

    # Also check for duplicates by title
    posts_with_title = client.get_posts_by_title(title)

    # Handle duplicates
    if cleanup_duplicates and len(posts_with_title) &gt; 1:
        print(f"  &#9888; Found {len(posts_with_title)} duplicate posts with title '{title}'")
        # Sort by updated_at to keep the most recent
        posts_sorted = sorted(posts_with_title, key=lambda p: p['updated_at'], reverse=True)
        post_to_keep = posts_sorted[0]
        posts_to_delete = posts_sorted[1:]

        print(f"    Keeping: {post_to_keep['slug']} (ID: {post_to_keep['id']}, updated: {post_to_keep['updated_at'][:10]})")

        for post in posts_to_delete:
            print(f"    Deleting: {post['slug']} (ID: {post['id']}, updated: {post['updated_at'][:10]})")
            success, error = client.delete_post(post['id'])
            if success:
                print(f"    &#10003; Deleted duplicate")
            else:
                print(f"    &#10007; Error deleting duplicate: {error}")

        # Use the kept post as the existing post
        existing_post = post_to_keep

    if existing_post:
        # Update existing post
        ghost_post['updated_at'] = existing_post['updated_at']
        success, error = client.update_post(existing_post['id'], ghost_post)
        if success:
            return 'updated'
        else:
            print(f"  &#10007; Error updating {slug}: {error}")
            return 'error'
    else:
        # Create new post
        success, error = client.create_post(ghost_post)
        if success:
            return 'created'
        else:
            print(f"  &#10007; Error creating {slug}: {error}")
            return 'error'


def main():
    """Main function to sync all posts."""
    import argparse

    parser = argparse.ArgumentParser(description='Sync Markdown posts to Ghost CMS')
    parser.add_argument('--dry-run', action='store_true', help='Preview changes without syncing')
    parser.add_argument('--limit', type=int, help='Limit number of posts to process')
    parser.add_argument('--post', type=str, help='Sync only a specific post (filename)')
    parser.add_argument('--cleanup-duplicates', action='store_true', help='Detect and remove duplicate posts with the same title')
    args = parser.parse_args()

    # Load environment variables
    load_dotenv(ENV_FILE)

    api_key = os.getenv('GHOST_ADMIN_API_KEY')
    if not api_key:
        print("Error: GHOST_ADMIN_API_KEY not found in .env file")
        sys.exit(1)

    # Initialize Ghost API client
    try:
        client = GhostAPIClient(GHOST_API_URL, api_key)
    except ValueError as e:
        print(f"Error: {e}")
        sys.exit(1)

    # Find all markdown files
    if args.post:
        # Sync specific post
        post_path = CONTENT_DIR / args.post
        if not post_path.exists():
            print(f"Error: Post not found: {post_path}")
            sys.exit(1)
        markdown_files = [post_path]
    else:
        # Sort in reverse order (newest first)
        markdown_files = sorted(CONTENT_DIR.glob('*.md'), reverse=True)
        if args.limit:
            markdown_files = markdown_files[:args.limit]

    total_posts = len(markdown_files)

    cleanup_msg = " (with duplicate cleanup)" if args.cleanup_duplicates else ""
    if args.dry_run:
        print(f"[DRY RUN MODE] Processing {total_posts} posts from {CONTENT_DIR}{cleanup_msg}...\n")
    else:
        print(f"Processing {total_posts} posts from {CONTENT_DIR}{cleanup_msg}...\n")

    # Track statistics
    stats = {
        'created': 0,
        'updated': 0,
        'error': 0,
        'skipped': 0
    }

    # Process each file
    for i, file_path in enumerate(markdown_files, 1):
        print(f"[{i}/{total_posts}] {file_path.name}")

        status = sync_post_to_ghost(client, file_path, dry_run=args.dry_run, cleanup_duplicates=args.cleanup_duplicates)
        stats[status] += 1

        if status == 'created':
            print(f"  &#10003; Created")
        elif status == 'updated':
            print(f"  &#10003; Updated")
        elif status == 'skipped' and not args.dry_run:
            print(f"  - Skipped")

        # Small delay to avoid rate limiting (skip in dry-run)
        if not args.dry_run:
            time.sleep(0.1)

    # Print summary
    print("\n" + "=" * 50)
    print("Sync Summary:")
    print(f"  Created: {stats['created']}")
    print(f"  Updated: {stats['updated']}")
    print(f"  Skipped: {stats['skipped']}")
    print(f"  Failed: {stats['error']}")
    print(f"  Total: {total_posts}")
    print("=" * 50)


if __name__ == '__main__':
    main()</code></pre><p>As you can see, this script</p><ul><li><p>fetches all Markdown files in a specific subdirectory</p></li><li><p>converts the source to HTML</p></li><li><p>uses the Admin API to create the post if not exists, or update existing posts</p></li><li><p>cleans up duplicate posts</p></li><li><p>provided keywords, tags, header images, social images, etc. in the expected format</p></li></ul><p>And while this is cool, do you know what's even cooler? Automating it, of course! So I set up a GitHub Actions workflow to run the script every time content changed:</p><pre><code>name: Sync Posts to Ghost CMS

on:
  push:
    branches:
      - main
    paths:
      - 'content/posts/**'
      - 'scripts/sync-to-ghost.py'
      - 'scripts/requirements.txt'
  workflow_dispatch:
    inputs:
      dry_run:
        description: 'Run in dry-run mode'
        required: false
        default: 'false'
        type: choice
        options:
          - 'true'
          - 'false'

jobs:
  sync-to-ghost:
    runs-on: ubuntu-latest
    permissions:
      contents: read

    steps:
      - name: Checkout repository
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: '3.11'
          cache: 'pip'

      - name: Install dependencies
        run: pip install -r scripts/requirements.txt

      - name: Sync posts to Ghost
        env:
          GHOST_ADMIN_API_KEY: ${{ secrets.GHOST_ADMIN_API_KEY }}
          PYTHONUNBUFFERED: 1
        run: |
          if [ "${{ github.event.inputs.dry_run }}" = "true" ]; then
            python scripts/sync-to-ghost.py --dry-run
          else
            python scripts/sync-to-ghost.py
          fi

      - name: Upload sync summary
        if: always()
        run: |
          echo "Sync completed. Check logs for details."</code></pre><p>This is great for getting all content uploaded <em>to</em> Ghost, but what about using Ghost to author new posts while still relying on Markdown as the single source of truth? Enter the reverse sync process!</p><h2>Syncing from Ghost to Markdown</h2><p>Fetching and uploading all Markdown content to Ghost is nice, but it would be even better if I didn't have to enter a code editor to author my blog posts. Don't get me wrong, I love NeoVim now, but writing a blog is not the same as writing code to me, so I would strongly prefer a lighter experience.</p><p>After uploading all my initial posts I quickly thought: Why can't I run the same process in reverse? If I can turn Markdown into HTML, the opposite shouldn't be rocket science either. So I set out to create a reverse syncing script.</p><pre><code>#!/usr/bin/env python3
"""
Ghost CMS Reverse Sync Script
Syncs posts from Ghost CMS to local Markdown files.
"""

import os
import re
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
from typing import Dict, List, Optional
from urllib.parse import urlparse

import jwt
import requests
import frontmatter
from markdownify import MarkdownConverter
from bs4 import BeautifulSoup
from dotenv import load_dotenv


# Configuration
GHOST_API_URL = "https://YOUR_DOMAIN_HERE"
CONTENT_DIR = Path(__file__).parent.parent / "content" / "posts"
STATIC_DIR = Path(__file__).parent.parent / "content" / "static" / "posts"
ENV_FILE = Path(__file__).parent.parent / ".env"


class GhostAPIClient:
    """Client for interacting with Ghost Admin API."""

    def __init__(self, api_url: str, admin_api_key: str):
        self.api_url = api_url.rstrip('/')
        self.admin_api_key = admin_api_key

        # Parse the admin API key
        try:
            self.key_id, self.key_secret = admin_api_key.split(':')
        except ValueError:
            raise ValueError("Invalid GHOST_ADMIN_API_KEY format. Expected 'ID:SECRET'")

    def _generate_jwt_token(self) -&gt; str:
        """Generate a JWT token for Ghost Admin API authentication."""
        iat = int(datetime.now(timezone.utc).timestamp())

        header = {
            'alg': 'HS256',
            'typ': 'JWT',
            'kid': self.key_id
        }

        payload = {
            'iat': iat,
            'exp': iat + 300,  # Token expires in 5 minutes
            'aud': '/admin/'
        }

        token = jwt.encode(payload, bytes.fromhex(self.key_secret), algorithm='HS256', headers=header)
        return token

    def get_all_posts(self) -&gt; List[Dict]:
        """Fetch all posts from Ghost with pagination."""
        all_posts = []
        page = 1

        while True:
            token = self._generate_jwt_token()
            url = f"{self.api_url}/ghost/api/admin/posts/"
            params = {
                'limit': 50,
                'page': page,
                'formats': 'html',
                'include': 'tags,codeinjection_head'
            }
            headers = {
                'Authorization': f'Ghost {token}',
                'Content-Type': 'application/json'
            }

            try:
                response = requests.get(url, params=params, headers=headers)
                response.raise_for_status()
            except Exception as e:
                raise Exception(f"Failed to fetch posts (page {page}): {e}")

            data = response.json()
            posts = data.get('posts', [])
            all_posts.extend(posts)

            # Check if there are more pages
            meta = data.get('meta', {}).get('pagination', {})
            if page &gt;= meta.get('pages', 1):
                break

            page += 1
            time.sleep(0.1)  # Small delay between requests

        return all_posts


class CustomMarkdownConverter(MarkdownConverter):
    """Custom converter that preserves code block languages."""

    def convert_pre(self, el, text, **kwargs):
        """Override pre tag conversion to preserve language classes."""
        if not text:
            return ''

        # Check if this pre contains a code element
        code = el.find('code')
        if code is not None:
            # Extract language from class attribute
            classes = code.get('class', [])
            if isinstance(classes, str):
                classes = classes.split()

            language = ''
            for cls in classes:
                if cls.startswith('language-'):
                    language = cls.replace('language-', '')
                    break

            # Get the code content
            code_text = code.get_text()

            # Return fenced code block with language
            # Ensure code_text ends with newline for proper fence formatting
            if not code_text.endswith('\n'):
                code_text += '\n'

            if language:
                return f'\n```{language}\n{code_text}```\n'
            else:
                return f'\n```\n{code_text}```\n'

        return super().convert_pre(el, text, **kwargs)


def convert_html_to_markdown(html: str) -&gt; str:
    """Convert Ghost HTML to Markdown, preserving code block languages."""
    return CustomMarkdownConverter(
        heading_style="ATX",
        bullets="-",
        escape_asterisks=False,
        escape_underscores=False,
        strip=['script', 'style']
    ).convert(html)


def download_images(html: str, slug: str, dry_run: bool = False) -&gt; None:
    """Download images from HTML to static directory."""
    soup = BeautifulSoup(html, 'html.parser')

    images_to_download = []
    for img in soup.find_all('img'):
        src = img.get('src')
        if not src or not src.startswith('https://static.brunoscheufler.com/posts/'):
            continue

        # Extract the full relative path from the URL
        # e.g., "https://static.brunoscheufler.com/posts/slug/gcloud/1.png"
        # -&gt; extract "slug/gcloud/1.png"
        parsed = urlparse(src)
        url_path = parsed.path  # e.g., "/posts/provisioning-k8s-cluster/gcloud/1.png"

        # Remove '/posts/' prefix to get relative path
        if url_path.startswith('/posts/'):
            relative_path = url_path[7:]  # Remove '/posts/' -&gt; "slug/gcloud/1.png"
        else:
            continue

        # Build target path preserving subdirectories
        target_path = STATIC_DIR / relative_path

        # Check if file exists
        if target_path.exists():
            print(f"    Image already exists: {relative_path}")
            continue

        images_to_download.append((src, target_path, relative_path))

    if not images_to_download:
        return

    if dry_run:
        # In dry-run mode, just log what would be downloaded
        for src, target_path, relative_path in images_to_download:
            print(f"    [DRY RUN] Would download: {relative_path}")
        return

    # Create parent directories and download
    for src, target_path, relative_path in images_to_download:
        try:
            target_path.parent.mkdir(parents=True, exist_ok=True)  # Create all parent dirs
            response = requests.get(src, timeout=10)
            response.raise_for_status()
            target_path.write_bytes(response.content)
            print(f"    Downloaded: {relative_path}")
        except Exception as e:
            print(f"    Warning: Failed to download {src}: {e}")


def extract_keywords_from_code_injection(codeinjection_head: Optional[str]) -&gt; List[str]:
    """Extract keywords from Ghost code injection HTML."""
    if not codeinjection_head:
        return []

    soup = BeautifulSoup(codeinjection_head, 'html.parser')
    keywords_tag = soup.find('meta', {'name': 'keywords'})

    if keywords_tag:
        content = keywords_tag.get('content', '')
        if content:
            return [k.strip() for k in content.split(',')]

    return []


def extract_ogimage_from_code_injection(codeinjection_head: Optional[str]) -&gt; Optional[str]:
    """Extract og:image URL from Ghost code injection HTML."""
    if not codeinjection_head:
        return None

    soup = BeautifulSoup(codeinjection_head, 'html.parser')
    ogimage_tag = soup.find('meta', {'property': 'og:image'})

    if ogimage_tag:
        content = ogimage_tag.get('content', '')
        if content:
            return content.strip()

    return None


def map_ghost_post_to_frontmatter(post: Dict) -&gt; Dict:
    """Map Ghost post fields to frontmatter."""
    # Extract date from published_at or created_at
    published_at = post.get('published_at') or post.get('created_at')

    # Map status
    is_published = post.get('status') == 'published'

    # Extract tag names
    tags = post.get('tags', [])
    topics = [tag['name'] for tag in tags if isinstance(tag, dict)]

    # Build path with /blog/ prefix
    slug = post['slug']
    path = f"/blog/{slug}"

    # Extract keywords from code injection
    keywords = extract_keywords_from_code_injection(post.get('codeinjection_head'))

    # Extract og:image from code injection
    ogimage_url = extract_ogimage_from_code_injection(post.get('codeinjection_head'))

    # Build frontmatter
    frontmatter_data = {
        'published': is_published,
        'type': 'BLOG',  # Default to BLOG, can be manually changed
        'path': path,
        'date': published_at,
        'title': post['title'],
        'keywords': keywords,  # Extracted from code injection, empty if not present
        'topics': topics
    }

    # Add headerImageUrl if feature_image exists
    if post.get('feature_image'):
        frontmatter_data['headerImageUrl'] = post['feature_image']

    # Add ogImageUrl if extracted from code injection
    if ogimage_url:
        frontmatter_data['ogImageUrl'] = ogimage_url

    return frontmatter_data


def generate_filename(post: Dict) -&gt; str:
    """Generate filename from Ghost post."""
    published_at = post.get('published_at') or post.get('created_at')
    date_str = published_at[:10]  # YYYY-MM-DD
    slug = post['slug']

    # If slug already has date prefix, use as-is
    if re.match(r'^\d{4}-\d{2}-\d{2}-', slug):
        return f"{slug}.md"

    return f"{date_str}-{slug}.md"


def sync_post_from_ghost(post: Dict, dry_run: bool = False) -&gt; str:
    """Sync a single post from Ghost. Returns status: 'created', 'updated', or 'skipped'."""
    filename = generate_filename(post)
    file_path = CONTENT_DIR / filename

    # Convert HTML to Markdown
    markdown_content = convert_html_to_markdown(post['html'])

    # Download images (or check which would be downloaded in dry-run)
    download_images(post['html'], post['slug'], dry_run=dry_run)

    # Map to frontmatter
    frontmatter_data = map_ghost_post_to_frontmatter(post)

    # Create frontmatter post
    new_post = frontmatter.Post(markdown_content, **frontmatter_data)

    if dry_run:
        print(f"  [DRY RUN] Would write: {filename}")
        return 'skipped'

    # Check if file exists
    status = 'updated' if file_path.exists() else 'created'

    # Write file
    with open(file_path, 'w', encoding='utf-8') as f:
        f.write(frontmatter.dumps(new_post))

    return status


def main():
    """Main function to sync all posts from Ghost."""
    import argparse

    parser = argparse.ArgumentParser(description='Sync posts from Ghost CMS to Markdown')
    parser.add_argument('--dry-run', action='store_true', help='Preview changes without syncing')
    args = parser.parse_args()

    # Load environment variables
    load_dotenv(ENV_FILE)

    api_key = os.getenv('GHOST_ADMIN_API_KEY')
    if not api_key:
        print("Error: GHOST_ADMIN_API_KEY not found in .env file")
        sys.exit(1)

    # Initialize Ghost API client
    try:
        client = GhostAPIClient(GHOST_API_URL, api_key)
    except ValueError as e:
        print(f"Error: {e}")
        sys.exit(1)

    # Fetch all posts from Ghost
    print("Fetching posts from Ghost...")
    try:
        posts = client.get_all_posts()
    except Exception as e:
        print(f"Error fetching posts: {e}")
        sys.exit(1)

    total_posts = len(posts)

    if args.dry_run:
        print(f"\n[DRY RUN MODE] Processing {total_posts} posts...\n")
    else:
        print(f"\nProcessing {total_posts} posts...\n")

    # Track statistics
    stats = {'created': 0, 'updated': 0, 'skipped': 0}

    # Process each post
    for i, post in enumerate(posts, 1):
        print(f"[{i}/{total_posts}] {post['title']}")

        status = sync_post_from_ghost(post, dry_run=args.dry_run)
        stats[status] += 1

        if status == 'created':
            print(f"  &#10003; Created")
        elif status == 'updated':
            print(f"  &#10003; Updated")

    # Print summary
    print("\n" + "=" * 50)
    print("Reverse Sync Summary:")
    print(f"  Created: {stats['created']}")
    print(f"  Updated: {stats['updated']}")
    print(f"  Skipped: {stats['skipped']}")
    print(f"  Total: {total_posts}")
    print("=" * 50)


if __name__ == '__main__':
    main()</code></pre><p>You can see this script does the following</p><ul><li><p>load all posts from the API</p></li><li><p>convert HTML to Markdown while respecting certain rules for code blocks, etc.</p></li><li><p>download images and media from posts into the expected static content directory</p></li></ul><p>I ran this a couple times to verify the format would be as close to the existing source as possible (having 230+ examples to choose from definitely helps with testing) and then ran a full reverse sync once.</p><p>And again, what's better than doing this manually? Running it on a schedule!</p><pre><code>name: Reverse Sync from Ghost CMS

on:
  schedule:
    # Run daily at 2 AM UTC
    - cron: '0 2 * * *'
  workflow_dispatch:
    inputs:
      dry_run:
        description: 'Run in dry-run mode'
        required: false
        default: 'false'
        type: choice
        options:
          - 'true'
          - 'false'

jobs:
  reverse-sync:
    runs-on: ubuntu-latest
    permissions:
      contents: write
      pull-requests: write

    steps:
      - name: Checkout repository
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: '3.11'
          cache: 'pip'

      - name: Install dependencies
        run: pip install -r scripts/requirements.txt

      - name: Run reverse sync
        env:
          GHOST_ADMIN_API_KEY: ${{ secrets.GHOST_ADMIN_API_KEY }}
          PYTHONUNBUFFERED: 1
        run: |
          if [ "${{ github.event.inputs.dry_run }}" = "true" ]; then
            python scripts/sync-from-ghost.py --dry-run
          else
            python scripts/sync-from-ghost.py
          fi

      - name: Check for changes
        id: check_changes
        run: |
          if [ -n "$(git status --porcelain)" ]; then
            echo "has_changes=true" &gt;&gt; $GITHUB_OUTPUT
          else
            echo "has_changes=false" &gt;&gt; $GITHUB_OUTPUT
          fi

      - name: Count changes
        if: steps.check_changes.outputs.has_changes == 'true'
        id: count_changes
        run: |
          # Count new files
          NEW_FILES=$(git status --porcelain | grep "^??" | wc -l | xargs)
          # Count modified files
          MODIFIED_FILES=$(git status --porcelain | grep "^ M" | wc -l | xargs)

          echo "new_files=$NEW_FILES" &gt;&gt; $GITHUB_OUTPUT
          echo "modified_files=$MODIFIED_FILES" &gt;&gt; $GITHUB_OUTPUT

      - name: Create Pull Request
        if: steps.check_changes.outputs.has_changes == 'true'
        env:
          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
        run: |
          # Configure git
          git config --global user.name "github-actions[bot]"
          git config --global user.email "41898282+github-actions[bot]@users.noreply.github.com"

          # Create branch with timestamp
          TIMESTAMP=$(date +%Y%m%d-%H%M%S)
          BRANCH_NAME="ghost-sync-$TIMESTAMP"

          # Create and switch to new branch
          git checkout -b "$BRANCH_NAME"

          # Add all changes
          git add content/posts/ content/static/posts/

          # Commit changes
          git commit -m "Sync posts from Ghost CMS

          - New posts: ${{ steps.count_changes.outputs.new_files }}
          - Updated posts: ${{ steps.count_changes.outputs.modified_files }}

          Automated sync performed at $(date -u +"%Y-%m-%d %H:%M:%S UTC")"

          # Push branch
          git push origin "$BRANCH_NAME"

          # Create PR with detailed body
          gh pr create \
            --title "Sync posts from Ghost CMS ($TIMESTAMP)" \
            --body "## Summary
          This PR contains changes synced from Ghost CMS.

          - **New posts**: ${{ steps.count_changes.outputs.new_files }}
          - **Updated posts**: ${{ steps.count_changes.outputs.modified_files }}
          - **Sync time**: $(date -u +"%Y-%m-%d %H:%M:%S UTC")

          ## Changes

          The following content was synced from Ghost:
          - Markdown post files in \`content/posts/\`
          - Static images in \`content/static/posts/\`

          ## Review Checklist

          - [ ] Review new post content and frontmatter
          - [ ] Check image downloads completed successfully
          - [ ] Verify frontmatter mapping (especially \`type\` and \`keywords\` fields)
          - [ ] Ensure no sensitive content was synced
          - [ ] Test builds locally if significant changes

          ## Notes

          This is an automated sync. The \`type\` field defaults to \`'BLOG'\` and may need manual adjustment for guides/tutorials. The \`keywords\` field is empty and may benefit from manual population." \
            --head "$BRANCH_NAME" \
            --base main

      - name: No changes detected
        if: steps.check_changes.outputs.has_changes == 'false'
        run: |
          echo "No changes detected. Ghost CMS is in sync with repository."</code></pre><div><hr></div><p>I hope this post serves as inspiration if you've ever thought about setting up something similar. I can totally recommend using Claude Code for the tedious work here, so you can focus on your content!</p>]]></content:encoded></item><item><title><![CDATA[On picking software for the long run]]></title><description><![CDATA[I've always enjoyed using good software.]]></description><link>https://www.brunoscheufler.com/p/2025-12-22-on-picking-software</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2025-12-22-on-picking-software</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Mon, 22 Dec 2025 00:08:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I've always enjoyed using good software. Convenient tools. The newest thing. Over time, I've come to realize that none of this matters if you don't own your data. Or if you have to worry that a company providing the platform you're on goes out of business. What matters first and foremost is ownership and independence, now more than ever.</p><p>I've had a journey over the last years of finding the values I care about most. In 2018, I <a href="https://brunoscheufler.com/2018-12-22-designing-my-new-portfolio/">chose to write my posts in Markdown</a> to ensure they'd stand the test of time. Back then, MDX was all the rage, but the more components you could add, the more components you would have to maintain. I chose to limit my content to the lowest common denominator every software could agree on. In 2020, I <a href="https://brunoscheufler.com/2020-10-20-rebuilding-my-portfolio-using-nextjs-tailwind/">moved to Next.js</a> as I saw Gatsby nearing its end of life. GraphQL was very exciting for some time, and a rich pipeline of plugins was promising, yet it fell victim to complexity by trying to solve all use cases. Earlier this year, I moved from Vercel to my own server. And yesterday, I moved this blog <a href="https://brunoscheufler.com/2025-12-22-bidirectional-markdown-syncing-for-ghost/">from Next.js to Ghost</a>, the last move in a series of steps that reinforced my commitment to owning my data and remaining independent.</p><p>Look, I don't think the world would end if I lost my content. It might be freeing to start from scratch, actually. But the latest string of vulnerabilities, as well as constant breaking changes convinced me it was time to settle on a boring yet reliable technology. I've learned the hard way that less is more. Maybe I'll host text files on a web server one day, and I'll be happy with it.</p><p>Another area I care about a lot is my collection of notes. I've been a happy Notion customer for five years. I admire the team's work in building one of the best products that helped me think and store my knowledge. And yet, I'm afraid of the day the company is acquired or the priorities change away from providing the best experience to seeking rent, as countless other companies have done in the past. <a href="https://obsidian.md/?ref=brunoscheufler.com">Obsidian</a> is an amazing alternative that may lack a feature or two but gets better every update. And the best part about it? Your notes are stored as Markdown, on your device.</p><p>We have the luxury of choosing the software we use. With every day that passes, it gets easier to choose an alternative that respects your choices. Luckily that doesn't mean you have to give up craftsmanship or polish anymore.</p>]]></content:encoded></item><item><title><![CDATA[Welcome to my blog!]]></title><description><![CDATA[I work at the intersection of software engineering and management in startups, scaling engineering orgs from 0 to 1.]]></description><link>https://www.brunoscheufler.com/p/about</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/about</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 21 Dec 2025 08:39:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p>I work at the intersection of software engineering and management in startups, scaling engineering orgs from 0 to 1.</p></blockquote><p>I write about my experience solving management and engineering problems in the context of startups. As I'm constantly learning new things along the way, topics will vary depending on what's most important to save for the future, as I'm creating a collection of resources of everything I think is relevant for building software and building and growing software companies.</p><p>That said, all timeless posts are tagged as <a href="__GHOST_URL__/?topic=Evergreen">Evergreen</a>, for your convenience. And if you're interested in my software engineering origin story, you can read about it <a href="__GHOST_URL__/blog/2023-09-03-looking-back-on-a-decade-in-software-engineering">here</a>.</p><h2>Personal Favorites</h2><ul><li><p><a href="__GHOST_URL__/blog/2022-12-31-whats-next">What's Next? (2022)</a></p></li><li><p><a href="__GHOST_URL__/blog/2022-09-18-making-architects-work-in-software-teams">Making Architects work in Software Teams (2022)</a></p></li><li><p><a href="__GHOST_URL__/blog/2022-09-04-steady-state-means-continuous-rewriting">Steady State means Continuous Rewriting (2022)</a></p></li><li><p><a href="__GHOST_URL__/blog/2020-09-01-fundamental-design-decisions-for-scalable-systems">Fundamental Design Decisions for Scalable Systems (2020)</a></p></li></ul><h2>At the intersection of engineering and management</h2><p>In recent years, I've worked as a software engineer, building backend services handling millions of requests per day, using technologies like Golang, GraphQL, PostgreSQL, TypeScript, and Node.js. I've also spent a fair share of my time building frontend web applications using React.js, Apollo, and XState. In addition to full-stack web engineering, I'm also building mobile applications for iOS using Swift and SwiftUI/UIKit. While I love trying out new and experimental technologies, I strongly prefer battle-tested solutions and boring technologies for production systems.</p><p>Over time, I helped scale engineering teams, onboarding new engineers, making architectural decisions, and helping build a culture of trust and ownership to enable engineers to do their best work. I worked together with software engineers, product managers, leadership, sales, and marketing, aligning engineering efforts with business goals and making sure we build the things our customers need most.</p><h2>What's my focus right now?</h2><p>Since <a href="__GHOST_URL__/blog/2024-04-23-joining-inngest">joining in May 2024</a>, I've been building distributed systems at <a href="https://www.inngest.com?ref=brunoscheufler.com">Inngest</a>, helping engineering teams to focus on creating value and reducing complexity.</p><h2>What do I value at work?</h2><p>I am a builder with strong action bias. I value ownership, trust, and great communication in teams. I think that the best teams listen closely to their customers and build things people want, at the cutting edge of technology. While tech is great, I strongly believe that humans are the most important part of any organization, and that empathy, trust, and communication are the most important skills for any team to succeed. I love sharing my knowledge and experience with others, and I am always looking to learn from others, no matter their background or experience.</p><p>I am not limited to engineering: I enjoy working with people, enabling teams, and solving problems, regardless of the domain. I have worked on product, engineering, design, sales, and marketing challenges, and I am always looking to learn more about how to build great products and great teams. In the end, I really care about building great products that solve unmet needs, and I am happy to help out wherever I can.</p><h2>Paying it forward</h2><p>I have been fortunate to have many great mentors and teachers in my career, and I am always looking to pay it forward. If you're thinking about founding or already building a startup or if you are working on a project that you'd like to get feedback on, feel free to reach out to me via email and I'll do my best to help you out or point you in the right direction.</p><h2>Colophon</h2><p>This site is built with <a href="https://nextjs.org/">Next.js</a> and <a href="https://tailwindcss.com/">TailwindCSS</a> and hosted on <a href="https://vercel.com">Vercel</a>. I am using <a href="https://fonts.google.com/specimen/Inter">Inter</a> as sans-serif typeface and <a href="https://fonts.google.com/specimen/Spectral">Spectral</a> as serif typeface. The accent color is #0200ff.</p>]]></content:encoded></item><item><title><![CDATA[Claude Code is the ChatGPT Moment for Software Engineering]]></title><description><![CDATA[Claude Code may be one of the most exciting product releases since ChatGPT.]]></description><link>https://www.brunoscheufler.com/p/2025-06-22-claude-code-is-the-chatgpt-moment-for-software-engineering</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2025-06-22-claude-code-is-the-chatgpt-moment-for-software-engineering</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 22 Jun 2025 12:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Claude Code may be one of the most exciting product releases since ChatGPT. It's been out since early February this year and I started using it a couple weeks ago, but it's consistently blowing my mind.</p><p>I had been planning a project to replace a critical data pipeline system currently hosted on Google Cloud with Kafka, and I could have built it in a couple days, but I really wanted to test what's possible with Claude Code. So, instead of trying to blindly prompt along, I sat down for an hour and wrote a spec outlining a high-level summary, a detailed rollout plan, system reliability considerations, metrics to track, a definition of done, and step-by-step tasks to follow. The result was more an engineering spec than a prompt.</p><p>Feeding all this context to Claude, the worst outcome could have been completely broken, incomplete, or outright bad code. Even so, I would have been better off than before, as the spec in itself was invaluable: Writing down all context and constraints for Claude to follow forced me to think clearly. Steering and scoping, not writing code, represents the actual value-add. This is one of the key themes I'll try to emphasize in this post.</p><p>Back to the refactor, Claude took an hour to author all changes, and to my surprise, the results were near perfect. The changes I did add by hand were not clearly described in the specification, so I couldn't even blame Claude. Yet, even with some ambiguity, Claude managed to produce a reasonable first iteration for a non-trivial problem. It followed instructions on breaking down the problem into smaller parts, as well as using established patterns found in similar implementations across the codebase to ensure consistency. It repeatedly fixed issues to end up with code changes that successfully compiled and passed linting errors.</p><p>I'm surprised by how well this worked, and it got me thinking.</p><h2>A short recap of recent history</h2><p>ChatGPT was first released in late November 2022. At that point, LLMs had been around for some time (GPT-1 was released in 2018, GPT-3 in 2020) but weren't accessible to consumers, neither technically (no public products) nor in terms of usability (models weren't fine-tuned to handle instructions). ChatGPT delivered a dead-simple chat interface with a model that reacted to instructions (GPT-3.5), and it mostly just worked.</p><p>The first iteration of instruction-tuned LLMs felt revolutionary for generating and editing text, as well as helping with learning, but the lack of first-class tools and integrations created lots of friction. Interestingly, <a href="https://arxiv.org/abs/2210.03629?ref=brunoscheufler.com">ReAct</a>, the first big paper on reasoning and tool use was released way before reasoning models and function calling were integrated into frontier models. In February 2023, the <a href="https://arxiv.org/abs/2302.04761?ref=brunoscheufler.com">Toolformer</a> paper outlined how to equip LLMs with tools including a calculator to help with basic arithmetic and access to search engines to ground knowledge and prevent hallucinations.</p><p>A couple months after ChatGPT's initial release, GPT-4 was introduced in early 2023, with function calling APIs <a href="https://openai.com/index/function-calling-and-other-api-updates/?ref=brunoscheufler.com">launching</a> in June that year.</p><p>Early LLM products relied heavily on semantic search and context compression techniques like RAG to find relevant knowledge without exceeding the painfully limited context windows. Due to limited or missing function calling capabilities, models were often built to return encoded instructions to calling systems for orchestrating further steps.</p><p>With the launch of Cursor in March 2023, a new generation of AI-enabled developer tools started making waves. Fast-forward to late 2024 and reasoning models were announced by most AI labs. Similar to the approach of <a href="https://arxiv.org/abs/2201.11903?ref=brunoscheufler.com">Chain-of-Thought (CoT)</a> prompting, reasoning models produce reasoning or <em>thought</em> tokens at inference-time, which are then used to generate the final response, effectively applying more compute time to come up with a response. OpenAI <a href="https://openai.com/o1/?ref=brunoscheufler.com">published</a> o1 in September 2024, Gemini 2.0 Flash Thinking <a href="https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/?ref=brunoscheufler.com">arrived</a> in December, and <a href="https://arxiv.org/abs/2501.12948?ref=brunoscheufler.com">DeepSeek-R1</a> was released in January 2025. Claude 3.7 Sonnet was <a href="https://www.anthropic.com/news/claude-3-7-sonnet?ref=brunoscheufler.com">released</a> in February 2025, featuring <a href="https://www.anthropic.com/research/visible-extended-thinking?ref=brunoscheufler.com">extended thinking</a> as a more fine-grained control over reasoning capabilities. This current iteration of models is <strong>significantly better</strong> at selecting the right tool for a problem, enabling more autonomous or <em>agentic</em> experiences.</p><p>Newer models are sufficiently more capable in most dimensions: Gemini 2.5 Pro features a context window of 1m tokens, compared to the 32k context window originally supported by GPT-4, the flagship model released two years earlier. This enables applications to include vastly more data, grounding the model and yielding more relevant results. Reasoning capability allows the latest models to solve more complex problems in various disciplines and interact with tools much more reliably and effectively.</p><p>In this context, the first research preview for Claude Code was <a href="https://www.anthropic.com/news/claude-3-7-sonnet?ref=brunoscheufler.com">released</a> on February 24, 2025, together with Claude 3.7 Sonnet. Just two months later, Anthropic announced the fourth generation of Claude models, and general availability for Claude Code.</p><h2>UX matters</h2><p>Claude Code embodies the state-of-the-art approach of equipping reasoning LLMs with tools to produce a smooth UX the same way early ChatGPT opened up LLMs to the public. While Cursor's Agent mode has been around since <a href="https://www.cursor.com/changelog/new-composer-ui-agent-commit-messages?ref=brunoscheufler.com">November 2024</a>, something about using Claude Code in a terminal feels more ergonomic than an IDE+Agent combination.</p><p>For one, Claude is remarkably good at picking the right tools while hiding unnecessary noise from lookup operations. The agent loop works near-autonomously, yet you can always interrupt the model and update instructions when you see it getting off-track. And lastly, the IDE integrations are the perfect balance between sharing context/using IDE features like diff views without being too imposing on the conversation flow.</p><p>I've been using Claude Code for projects I've always wanted to do, yet never got around to doing. Tedious tasks like writing CI/CD pipelines or Infrastructure as Code or upgrading and migrating dependencies essentially disappear, so you don't waste time on tasks that don't add business value. Besides, having multiple tabs of Claude Code running in parallel feels like managing a small team, even if you can't actually step away (yet).</p><p>As with any new tool, there are rough edges. Resuming and forking conversations works out of the box, but with long threads I do fear that previous context influences subsequent responses too much, steering the model in the wrong direction. Restarting conversations from scratch isn't great either, so using project-level memory and docs like <a href="https://brunoscheufler.com/blog/2020-07-04-documenting-design-decisions-using-rfcs-and-adrs">RFCs and ADRs</a> may be a good compromise.</p><p>In the long run, I believe Claude Code will augment most internal tooling. There are tons of MCP Servers readily available and integrating the tools you use every day (issue tracking, observability) is pretty much a done deal. As models become cheaper and more powerful over time (better reasoning, even larger context windows, faster response times), tools like Claude Code will be able to write (author, test, review, maintain) better code, more autonomously.</p><p>If this sounds scary, let me bring up two important points.</p><p>Claude Code is incredibly good at writing code but your results will be subpar if you're not steering it well. Just as you wouldn't let an engineer work on a codebase without objectives, check-ins, or reviews, Claude can't read your mind (yet). While you may spend less time writing code with Claude, you will spend more time thinking about the code that needs to be written. And why. And when. Tools have always existed, and this certainly is the best toolbox on the market. Yet, what really matters are the decisions you make and the outcomes you achieve.</p><p>Second, real business value isn't in wrangling dependencies, figuring out how Terraform works, setting up an S3 bucket, writing GitHub Actions workflows, fixing linting issues, or refactoring a codebase. Real business value comes from interactions between people. It's created in understanding customer needs and building relationships. Real value is created by building teams and aligning people on shared goals. With tools like Claude Code, we have a lot more time for that. What a time to be a software engineer!</p><h2>A note of caution</h2><p>Claude Code and LLMs in general are very powerful tools and can solve writing tasks better than most people in a fraction of the time. They can support knowledge discovery and the learning process, but I strongly believe that you should not use them to avoid friction and difficulty in the first place, even if it's tempting, even if the results will be just as good or much better than doing it yourself. Here's why: Asking Claude to explain concepts like a codebase works incredibly well, but personally, I learn a lot from following a system end-to-end. While there are many kinds of people out there, I'm sure the following may sound familiar.</p><p>When I started writing my first pieces of software more than a decade ago, I spent hours and hours of my time producing bugs, and trial-and-erroring my way to a working piece of code. And while the resulting projects may not have been a big commercial success, they helped me learn. Every time I stumbled over some issue, I learned. Friction triggers awareness, which then drives the learning process.</p><p>Having an LLM apply learning tools like Socratic questioning ensures you don't accidentally let AI do all the work, which can be great. Yet, I wonder what academia and vocational training for computer science/software engineering will look like in a world where tools like Cursor and Claude Code are commonplace. Are people truly going to pick up concepts? Did people have the same questions when the internet and search engines came up? (Yes!)</p><p>Another risk of excessive AI use is that you have LLMs replace your truly important writing. I'm not talking about drafting some official-sounding letters, outreach message, website copy, or shitpost for Twitter. As engineers, we produce a lot of written content, both internally and externally. Writing helps me think clearly. Translating a problem in your head to a piece of paper in front of you forces you to communicate in a common language. You can use LLMs to proofread written content, but for the first draft, it's your turn.</p><p>To repeat, your value (as a software engineer specifically, and most related disciplines) is in identifying valuable problems to solve, weighing opportunity cost and deciding which problems to solve first, making necessary tradeoffs to scope problems into a realistic schedule, and delivering a solution. Just as a team lead delegates tasks, you're not measured by the lines of code you write but the outcome. So in a way, AI levels everyone up to a team lead of AI workers.</p><p>In this world, communication skills are more important than ever. Specifically, clear and concise writing. In addition, curiosity and willingness to try out new things. If you're not ready to learn and make some mistakes, it's infinitely harder to discover even better ways of working.</p><h2>Wrapping Up</h2><p>Claude Code, Cursor, and other AI products are very impressive and useful. While I can't gauge the full impact of progressively shifting to AI-driven software engineering, I can see the change in perception in my own ways of working. I can't wait to see how this space keeps evolving!</p>]]></content:encoded></item><item><title><![CDATA[Looking Back on 2024]]></title><description><![CDATA[Another year is coming to a close, so it&#8217;s time for another annual review!]]></description><link>https://www.brunoscheufler.com/p/2024-12-30-looking-back-on-2024</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-12-30-looking-back-on-2024</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Mon, 30 Dec 2024 18:00:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Another year is coming to a close, so it&#8217;s time for another annual review! This has become a personal tradition, check out <a href="https://brunoscheufler.com/blog/2023-12-31-looking-back-on-2023">last year&#8217;s review</a> if you haven&#8217;t done so already!</p><p>Early in 2024, I was working hard on building CodeTrail, a new take on documentation for engineering teams. During this time, I learned a lot about sales, marketing, and distribution. While Tim was working hard on building the product, I took over sales, creating cold outreach sequences, calling up potential customers, running demos, following up to leads &#8212; to make it short, I did everything under the sun to get our first customers on board. Unfortunately, due to a number of factors, we were bound to fail.</p><p>Neither me nor Tim had a large platform, both of us previously worked on high-growth startups but we underestimated the difficulty of activating our networks to discover the best pilot customers. We were solving a problem both of us experienced in our previous jobs, but in a market of layoffs and fears of an economic downturn, onboarding software understandably wasn&#8217;t on people&#8217;s minds. While we were able to craft a high-quality product, we were fighting an uphill battle in one of the most risk-averse markets on the planet.</p><p>After half a year, Tim and I decided to wind down our efforts on CodeTrail in order to reflect what we were missing to succeed and plan the next steps accordingly, removing any barriers to retry a couple years in the future.</p><p>While Tim is focusing on academia, I decided it made most sense to join a fast-moving team I could grow with. I spent March applying at and interviewing with 30+ companies in different growth stages and industries. At one point in my discovery stage, I got introduced to <a href="https://www.inngest.com/?ref=brunoscheufler.com">Inngest</a>, a Bay Area/US startup solving the hard problem of building reliable software, a passion area of mine I spent the past years of my life working on. After the first interviews with the team, I was fully convinced that this was the best possible option out there.</p><p>I joined the engineering team at Inngest exactly 8 months ago in late April. Since then, we&#8217;ve increased our product usage nearly 20x, which introduced scaling challenges I had the opportunity to solve.</p><p>During my first weeks of onboarding, I built batch keys, a new feature to group event batches by a user-supplied expression. This is helpful in multi-tenant scenarios, for example to group in-app notifications for a given user, or creating massive mailing campaigns.</p><p>After my first project, I started focusing on systems and infrastructure, joining our founding engineer Jack Williams in implementing a critical new service for disaggregating access to our primary data store, and implementing a sharding strategy for horizontal scaling. These improvements enabled us to handle 100k+ QPS comfortably without a second of downtime. If you&#8217;re curious about the full story, you can read more on the <a href="https://www.inngest.com/blog/sharding-at-inngest?ref=brunoscheufler.com">Inngest blog</a>.</p><p>After ensuring our infrastructure could handle the increased demand, I moved on to the core queueing systems, implementing a new architecture to improve multi-tenant fairness and throughput globally and for each account. Once this was rolled out, I spearheaded sharding efforts around the queue, unlocking horizontal scaling to handle a 5x increase in queue throughput.</p><p>I&#8217;ve learned a lot about designing, building, and operating distributed systems, performing gradual rollouts without degrading system availability, and engineering in high-growth environments. For the last months of 2024, I&#8217;ve been working on a major upcoming product and infrastructure feature, unlocking enterprise security, higher end-to-end throughput at lower latency, and a new execution model altogether. I&#8217;m excited to share more news on this soon!</p><p>Joining a fully-remote team also gave me the opportunity to meet everyone in person in Lisbon and San Francisco for company-wide and engineering offsites.</p><p>If you had told me what was about to happen throughout the year in January, I wouldn&#8217;t have believed you.</p><p>I&#8217;m forever grateful for the warm welcome by the Inngest team. I&#8217;m having an absolute blast working together with everyone and learning new things every day. Joining Inngest reignited my passion of engineering highly-scalable systems, making trade-offs to fit the growth stage, and building with a strong vision in mind.</p><p>Moving on to 2025, I&#8217;m excited to share big personal news relatively soon. <em>Onward and upward.</em></p><p>&#8212; Bruno</p>]]></content:encoded></item><item><title><![CDATA[Understanding When to Use Redis]]></title><description><![CDATA[Before joining Inngest earlier this year, I had never used Redis (or its successor, Valkey) or any comparable system.]]></description><link>https://www.brunoscheufler.com/p/2024-11-09-understanding-when-to-use-redis</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-11-09-understanding-when-to-use-redis</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sat, 09 Nov 2024 11:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Before joining <a href="https://www.inngest.com/?ref=bruno-redis-intro-post">Inngest</a> earlier this year, I had never used Redis (or its successor, Valkey) or any comparable system. Whenever I planned data persistence for a new service architecture, I would choose a reliable data store like Postgres. I would argue that ACID guarantees and a track record of stability outweigh any shiny technologies promising unlimited scalability or complete flexibility.</p><p>It&#8217;s awkward to admit, but I actually believed Redis was just a key/value store. More specifically, I never imagined keys could be anything other than bytes, let alone powerful data structures. The second I started to read up on Redis, I changed my mind. After spending most of my working hours the past six months working with Redis, I want to present cases for and against using Redis, Valkey, or any related technology.</p><p>Before we get into this, let&#8217;s consider why we need to persist data, why it sometimes makes sense to choose a data store optimized for specific access patterns, and why Redis is a great match if you&#8217;re looking for a high-performance data structure store, both persistent and ephemeral.</p><h2>Applications are rarely stateless</h2><p>This probably won&#8217;t come as a shock, but it&#8217;s good to When you&#8217;re building a SaaS product or any piece of modern software really, you&#8217;ll need to persist state. This may be anything from data required for operations over user data all the way to data for bookkeeping like logs and history.</p><p>In old times, this state could have lived on a single machine, but nowadays, you&#8217;ll have to consider fault tolerance and chances are, you need to serve customers in multiple regions around the world. Simply put, one server storing all state in memory doesn&#8217;t cut it anymore.</p><p>There are many data stores, offering different durability guarantees, data modeling paradigms, scaling options, and other properties. And this is where it gets interesting: How do you choose a data store suitable for your application?</p><h2>Data structures vs. relations</h2><p>I grew up in a world of relational database management systems. Tables became my hammer for every problem. This is obviously silly: Relational databases have their use cases, but forcing data into a model it wasn&#8217;t made for will cause problems down the road.</p><p>Most notably, this means poor ergonomics and performance issues: Traditional relational databases are optimized for index-heavy access operations, carefully scanning as few rows as possible. Complex query planners are designed to optimize data access based on historical statistics. You can select index types, sure, but a lot of performance characteristics are predetermined and leave little control.</p><p>Redis isn&#8217;t opinionated on how you structure your data. Instead, it provides a wide range of in-memory data structures: Strings, Hashes (maps, dictionaries), Lists, Sets, Sorted Sets (my favorite), and more.</p><p>You might wonder why that&#8217;s exciting, aren&#8217;t those just the basic building blocks of any programming language? You&#8217;d be right and wrong: When you&#8217;re building applications for a single node, these data structures are omnipresent. Redis allows you to keep this simple data layout while allowing you to access and modify this state from any client, written in any language.</p><p>As simple as these data structures sound, they cover almost any use case: Want to implement a queue? Use a list or, if you&#8217;re fancy, a sorted set. Want to index data for exact matches? Use a hash for O(1) access. Locks? You&#8217;re <a href="https://redis.io/docs/latest/develop/use/patterns/distributed-locks/?ref=brunoscheufler.com">covered</a>.</p><p>If this sounds exciting, you may wonder if there are any downsides. That depends on your requirements. Like most decisions in engineering, there&#8217;s no one good or bad solution out there.</p><h2>Choosing the right data store</h2><p>Engineering is all about tradeoffs, and this applies to selecting the right data store. If we&#8217;re building a latency-sensitive production system, we can lay out some hard requirements:</p><ul><li><p><strong>Availability</strong>: The system must withstand production workloads without downtime. This may require failover instances, multi-zone or multi-region replication, and most definitely</p></li><li><p><strong>Durability</strong>: Data must be persisted on disk. If the machine unexpectedly goes down, every stored record must be present when it becomes available again. To prevent losing data from missing writes during downtimes, see the Availability section for measures like automatic failover.</p></li><li><p><strong>Horizontal Scaling</strong>: Capacity planning is hard, and relying on a single instance can only buy so much time. Zero-downtime instance resizing is often impossible. Adding new nodes to the system must be easy and downtime during potential rebalancing operations must be kept to a minimum.</p></li></ul><p>There are other requirements including reliable disaster recovery, the remaining ACID properties (atomicity, consistency, isolation).</p><p>Redis is interesting in that it offers atomic operations on a global keyspace, supports durability with snapshots and append-only files, and allows horizontal scaling through sharding the global keyspace across a distributed cluster. You can even write Lua scripts to support atomic operations across multiple Redis commands.</p><p>Yet, there are downsides: <strong>Redis is, at its core, single-threaded</strong>. And even though individual operations are really fast, you&#8217;ll inevitably run into CPU bottlenecks if you&#8217;re unable to distribute your data across shards. Horizontal scaling only works through sharding, but once you distribute your keys across different slots, you lose atomicity: Multi-slot operations are prohibited, and that includes Lua scripts.</p><p>I&#8217;m really excited about the release of Valkey 8: The Valkey maintainers have put in lots of effort to implement <a href="https://valkey.io/blog/unlock-one-million-rps/?ref=brunoscheufler.com">I/O multithreading</a> while keeping execution single-threaded and sequential in nature, which boosts throughput and reduces latency across the board.</p><div><hr></div><p>Redis and its derivatives are incredibly versatile data stores enabling applications both small and large in scale. If you&#8217;re aware of the pitfalls when it comes to scaling using clustering and sharding, you can get really far with this one component.</p>]]></content:encoded></item><item><title><![CDATA[The First 120 Days at Inngest]]></title><description><![CDATA[Hey!]]></description><link>https://www.brunoscheufler.com/p/2024-08-31-the-first-120-days-at-inngest</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-08-31-the-first-120-days-at-inngest</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sat, 31 Aug 2024 11:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey! It&#8217;s been pretty quiet around here for a while &#8212; more than I&#8217;d like but it&#8217;s been a busy couple of months since I joined <a href="https://www.inngest.com/?ref=bruno-120-days-post">Inngest</a> earlier this year in April. I&#8217;m completing some personal projects too at the moment and trying to get back to a more regular publishing cadence from September on. In the meantime, enjoy this blog post!</p><h2>The search for a new challenge</h2><p>After multiple attempts at building my own company, I decided it was time to switch things up. I was looking for a talented team focused on building a world-class product, preferably in the infra or DevTools space, preferably in the Bay Area. Around March, the landscape for engineering jobs wasn&#8217;t ideal, if you were aiming to get into BigTech. Fortunately, my favorite company environments start around product-market fit and scaling up.</p><p>Once I started interviewing with Darwin and the folks at Inngest, I immediately noticed the high energy and conviction with which they&#8217;re building a remarkable product. I ran some due diligence and checked on the competition, and even four months later our product is still the most powerful and easiest to learn out there. I was hooked, and luckily the team felt the same way, so I joined at the end of April.</p><h2>The first months at Inngest</h2><p>When I joined, I did multiple pairing sessions with Darwin and Tony to grasp the fundamentals of the engineering setup. In my second week, I started working on a critical project with Jack, the details of which have justified a <a href="https://www.inngest.com/blog/sharding-at-inngest?ref=brunoscheufler.com">blog post</a> on the official company blog. After launching a set of infrastructure improvements to pool Redis connections in a dedicated service, we built, tested, and launched a sharding strategy within weeks. This was incredibly exciting and important for the infrastructure at the same time.</p><p>Fast forward to today, I&#8217;ve fully onboarded, worked on the core infrastructure, and already rolled out several mission-critical improvements which I&#8217;m very proud of. I&#8217;m also extremely grateful to everyone at Inngest for helping out wherever possible and sharing knowledge.</p><p>I&#8217;m writing this post on the plane back from Lisbon, where I just concluded my first offsite with the team. Working remotely can be hard because social connections and team bonding are hard to replicate in virtual meeting rooms. Having a week to get to know each other, tell stories, discuss current issues, and align everyone on the greater vision makes up for this. And it was an absolute blast.</p><h2>The challenges ahead</h2><p>Inngest, the product, has grown in usage by multiple orders of magnitude since I joined. Naturally, this presents a couple of challenges to the infrastructure we&#8217;ve been running for a while now. We&#8217;ve been hard at work designing and building the next version of many core infrastructure components to enable seamless growth for the next couple of years. We&#8217;ve made several key decisions that I&#8217;d love to share, but it will take some more time until we can unveil what we&#8217;re building.</p><p>At the same time, we&#8217;ll have to grow the team to free up capacity for building product features and other important work. We&#8217;re still in a stage where we operate with incredibly high autonomy, and taking ownership of areas of responsibility is expected.</p><p>I truly can&#8217;t wait for the months ahead!</p>]]></content:encoded></item><item><title><![CDATA[Enhancing Scalability and Reducing Latency Without Missing a Beat]]></title><description><![CDATA[At Inngest, we serve hundreds of millions of events per day.]]></description><link>https://www.brunoscheufler.com/p/2024-07-04-enhancing-scalability-and-reducing-latency-without-missing-a-beat</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-07-04-enhancing-scalability-and-reducing-latency-without-missing-a-beat</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Thu, 04 Jul 2024 11:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At <a href="https://www.inngest.com/?ref=bruno-state-coordinator-post">Inngest</a>, we serve hundreds of millions of events per day. Each event can trigger a function run comprising multiple steps. Users can return blobs and other data from these steps, necessitating incredibly fast reads and writes to our data store to keep latency as low as possible. Function runs can be delayed up to many months, so this data needs to be persisted over an arbitrary amount of time.</p><p>This immense volume of data is stored in our primary Redis-based data store, presenting significant challenges in terms of throughput and latency due to Redis&#8217; single-threaded design. Recently, our main cluster approached critical capacity limits, with peak CPU utilization and dwindling storage space threatening our system&#8217;s stability and performance.</p><p>Over the past six weeks, we implemented a series of infrastructure improvements that significantly reduced latency and improved long-term scalability. Impressively, we achieved this without a millisecond of downtime. In this post, I&#8217;ll give you an overview of the changes we&#8217;ve been working on, and focus on a new proxy service in the critical path.</p><p>To avoid degrading end-to-end latency and prevent storage capacity issues, we implemented a multi-stage plan focused on enhancing Inngest&#8217;s core components. Our first step was to <a href="https://brunoscheufler.com/blog/2024-06-29-harnessing-the-power-of-go-interfaces">introduce a unified state interface</a>, consolidating all state operations in the critical path in a single place. This way, improvements are propagated across the entire system.</p><p>Next, we introduced a coordination service to decouple our services from direct connections to the primary data store. Instead, requests are routed through a lightweight proxy, which supports sharding and caching for better efficiency. Long-term, this architecture allows us to replace the current state store with a completely new backend optimized for Inngest&#8217;s access patterns and SLOs.</p><p>Let&#8217;s delve into the implementation of the state coordinator.</p><h2>Implementing the State Coordinator</h2><p>After defining the new <a href="https://github.com/inngest/inngest/blob/3a2ea1b4c0798cb1740cdc0c3bda7abdb32e43f3/pkg/execution/state/v2/interfaces.go?ref=brunoscheufler.com#L28">state interface</a>, we immediately started work on the new coordination service, routing all state interface requests to the current or future state store. We designed the interfaces with the new service in mind, so the handover was seamless.</p><p>Accessing the state store is a critical operation in our system because every function run involves reading and writing events, step outputs, and other metadata. Failure in these operations could result in significant performance degradation across the system.</p><p>We designed the state coordinator to offer high throughput at the lowest possible latency. We implemented resilient error handling by retrying transient errors introduced by the network roundtrip, as well as introducing observability on both the client and server side, giving us insights into service performance and potential issues as early as possible.</p><p>By having all services connect to very few state coordinator instances, Redis connections are pooled in a single place, making it easy to understand bottlenecks and scale the system accordingly.</p><p>The state coordinator was built with robust sharding and caching mechanisms, utilizing consistent or highest random weight (HRW) hashing. This approach ensures uniform data distribution and high cache hit ratios, by routing requests to instances most likely to store the data in question. Sharding can be achieved by applying the same logic on the client to determine which group of state coordinators serving a single state store cluster should be invoked. Adding a new shard is as easy as provisioning another cluster and gracefully reconfiguring existing services.</p><h2>Rolling out</h2><p>To ensure a smooth rollout, we devised a strategy to gradually enroll Inngest customers in using the state coordinator. We leveraged the state interface to create a rollout layer using <a href="https://brunoscheufler.com/blog/2024-05-19-derisking-product-rollouts-with-feature-flags">feature flags</a>, enabling controlled, percentage-based rollouts across various customer segments: Internal users, free users, paid customers, and finally, enterprise customers.</p><p>To mitigate the risk of unexpected errors, we implemented robust fallback logic, ensuring any failed requests are retried using the previous implementation. In the worst case, the customer wouldn&#8217;t even notice that our new service failed, and no data would be lost in the process.</p><p>These safeguards gave us the confidence to roll out the new code to production very quickly, as we could flip a switch and fall back within a second. After thoroughly testing the new service internally, we enrolled all non-enterprise customers within a week. During the entire rollout, we tracked rollout progress combined with operation success and latency metrics using a custom dashboard, which made it easy to catch any problems, such as slow or failing requests. In every way, the rollout started smoothly.</p><p>After continuing to evaluate our system performance and allocating more compute resources to proactively rule out bottlenecks, we enrolled all enterprise customers without a hitch.</p><p>The state coordinator has been running in production for over a month now. On average, we are serving 300k rpm, with a p99 of less than 20ms, including network roundtrip latency. This is a fraction of the traffic going to our primary data store, yet it accounts for the largest response sizes due to user-generated state outputs.</p><h2>Post-Rollout</h2><p>Through this project, I gained valuable insights into building and deploying critical infrastructure, designing for gradual customer enrollment, creating effective dashboards for monitoring system health, and effectively communicating project updates with the team.</p><p>After conducting the rollout and running the state coordinator in production, there are a couple improvements I&#8217;d look into the next time we run a similar project. First, percentage-based rollouts are an amazing tool to slowly ramp up usage to spot bottlenecks early. Simply flipping a switch and routing all production traffic to a new system without safeguards would be reckless at best.</p><p>Percentage-based rollouts get more predictable the more uniformly distributed your enrollment key is. This is similar to sharding where the wrong sharding key can cause a skewed data distribution. If you roll out based on account IDs, it&#8217;s fairly likely that you encounter a long tail of accounts with little to no usage, with a few accounts making up most of the activity in your system. All account IDs look the same to a rollout system, so it&#8217;s possible that you start seeing a huge spike after increasing your percentage by a microscopic amount.</p><p>Knowing your product usage metrics is crucial for understanding how choosing a certain enrollment key will affect the system. One solution for this problem could be to create multiple segments based on different historical usage and adjust the entities in the buckets over time, leading to higher predictability.</p><div><hr></div><p>This has been my first major project to work on over at <a href="https://www.inngest.com/?ref=bruno-state-coordinator-post">Inngest</a>, and it&#8217;s been incredibly rewarding. Thanks to Jack, Darwin, Tony, and all the other folks on the team for helping out and making this possible!</p>]]></content:encoded></item><item><title><![CDATA[Restoring Go Errors in gRPC Services]]></title><description><![CDATA[In my last post, I talked about the benefits of using interfaces in Go.]]></description><link>https://www.brunoscheufler.com/p/2024-06-30-restoring-go-errors-in-grpc</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-06-30-restoring-go-errors-in-grpc</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 30 Jun 2024 11:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In my <a href="https://brunoscheufler.com/blog/2024-06-29-harnessing-the-power-of-go-interfaces">last post</a>, I talked about the benefits of using interfaces in Go. In short, interfaces are a powerful tool to specify the behavior of objects. However, Go interfaces lack a mechanism to specify a predefined set of errors that can be returned to narrow the definition, such as using a <code>throws</code> keyword like in <a href="https://docs.swift.org/swift-book/documentation/the-swift-programming-language/errorhandling/?ref=brunoscheufler.com#Specifying-the-Error-Type">Swift</a> or Java. An error returned by an interface could be any error. This is especially interesting when one implementation involves a network roundtrip that could introduce different errors than a location implementation.</p><p>When refactoring a production codebase, you don&#8217;t always have the budget or desire to rewrite the code consuming an interface. New implementations must match the previous logic down to the returned errors. Let&#8217;s see how this can be done when using gRPC.</p><h3>Errors in Go</h3><p>Go is built on conventions, including for handling errors. In Go, functions return errors alongside values. Errors can be wrapped and returned again, forming a chain of errors. This is quite nice as it helps you trace what went wrong, and where.</p><p>Below is an example demonstrating error chaining:</p><pre><code>func getOneRow(sql string, args []any) (any, error) {
&#9;res, err := db.Do(sql, args)
&#9;if err != nil {
&#9;&#9;return nil, fmt.Errorf("could not run command: %w", err)
&#9;}
&#9;return res, nil
}

func loadUser(userId string) (User, error) {
&#9;user, err := accessDb("select * from user where userId = $1;", userId)
&#9;if err != nil {
&#9;&#9;return nil, fmt.Errorf("could not access database: %w", err)
&#9;}
&#9;return user, nil
}

func isAdmin(userId string) (bool, error) {
&#9;user, err := loadUser
&#9;if err != nil {
&#9;&#9;return false, fmt.Errorf("could not load user: %w", err)
&#9;}
&#9;return user.isAdmin, nil
}</code></pre><p>In some layer of the call hierarchy, internal errors are converted into business logic errors: Not finding a row in a database may be relevant for internal code, but a user cares about the fact that their account doesn&#8217;t exist because they&#8217;ve signed up with a different email address. User-facing errors can be added into the error chain and retrieved using <code>errors.Is</code>.</p><h2>Errors over the network</h2><p>All these conventions work as long as your code runs in a single process. When sending requests over a network using libraries like gRPC, traditional error handling conventions can be challenging due to differing error serialization methods. gRPC offers error codes and details for attaching rich information to errors, which is great when all consumers expect gRPC-style error handling. Sometimes, error chains are just converted to their string representation. Good luck checking a string for the existence of a specific error.</p><p>Depending on your use case, the fact that functions are invoked over the network shouldn&#8217;t be exposed to the downstream consumers of your APIs, though. When defining an interface like <code>ImageCreator</code> below, you specify expected behavior without detailing network-specific errors or concrete error types.</p><pre><code>type ImageCreator interface {
&#9;CreateImage(data CreateData) (Image, error)
}

type CreateData struct {
&#9;Prompt string
}

type Image struct {
&#9;URL url.URL
}</code></pre><p>In other words, you can&#8217;t expect code consuming the <code>ImageCreator</code> interface to handle gRPC-specific errors.</p><h2>Restoring Errors</h2><p>Imagine we&#8217;re using the <code>ImageCreator</code> service outlined above to generate a header image for blog posts (how original).</p><pre><code>func generateHeaderImage(creator ImageCreator) (url.URL, error) {
&#9;img, err := creator.CreateImage(CreateData{"abstract blog post header"})
&#9;if err != nil {
&#9;&#9;if errors.Is(err, ErrInvalidPrompt) {
&#9;&#9;&#9;showInvalidPromptDialog()
&#9;&#9;&#9;return nil, nil
&#9;&#9;}

&#9;&#9;return nil, fmt.Errorf("something went wrong: %w", err)
&#9;}

&#9;return img.URL, nil
}</code></pre><p>To inform a user about passing an invalid prompt, we define the following error</p><pre><code>var ErrInvalidPrompt  = fmt.Errorf("invalid prompt")</code></pre><p>Consider two concrete implementations: A local service running in the same process, and a remote version running on beefier machines, communicating via gRPC.</p><p>Both implementations can return this error, but it might be hidden in a chain of errors</p><pre><code>could not create image: invalid prompt
could not send create image request: could not create image: invalid prompt</code></pre><p>Both of these are completely valid, and we want to surface them to the user with a proper error dialog! By using <code>errors.Is()</code> in the client implementation, we can traverse the error chain until we find the quota exceeded error. For the server, it won&#8217;t be this easy. Remember that most implementations simply return a string! So what can we do here?</p><p>Most implementations use error codes for retaining application-specific context. A strategy involves detecting user-facing errors before returning them, transforming them into a network representation and back to a concrete error on the client.</p><p>Below is an example implementation of this strategy.</p><h3>Common</h3><pre><code>var userFacingErrors := []error{
&#9;&#9;ErrQuotaExceeded,
}</code></pre><p>First, we define a slice of known errors. You could add any user-facing error here.</p><h3>Server</h3><pre><code>func preProcessError(err error) error {
&#9;for _, ufe := range userFacingErrors {
&#9;&#9;if errors.Is(err, ufe) {
&#9;&#9;&#9;return ufe
&#9;&#9;}
&#9;}

&#9;return err
}

// CreateImage implements imagecreator.CreateImage
func (s *server) CreateImage(ctx context.Context, in *pb.CreateImageRequest) (*pb.CreateImageReply, error) {
&#9;img, err := createImageOnServer(in)
&#9;if err != nil {
&#9;&#9;return nil, status.New(codes.InvalidArgument, preProcessError(err))
&#9;}

&#9;return &amp;pb.CreateImageReply{Image: img}, nil
}</code></pre><p><code>preProcessError</code> will iterate over a range of all known user-facing errors and, if included in the given error, return only the known error. This will shorten error chains and return only a single known error.</p><h3>Client</h3><pre><code>func restoreError(reason string) error {
&#9;for _, err := range userFacingErrors {
&#9;&#9;if err.Error() == reason {
&#9;&#9;&#9;return err
&#9;&#9;}
&#9;}

&#9;return nil
}

type imageCreatorClient {}

func (i *imageCreatorClient) CreateImage(data CreateData) (Image, error) {
&#9;c := pb.NewGreeterClient(conn)

&#9;ctx, cancel := context.WithTimeout(context.Background(), time.Second)
&#9;defer cancel()
&#9;r, err := c.CreateImage(ctx, &amp;pb.CreateImageRequest{data})
&#9;if err != nil {
&#9;&#9;s := status.Convert(err)
&#9;&#9;if restored := restoreError(s.Message()); restored != nil {
&#9;&#9;&#9;return Image{}, restored
&#9;&#9;}

&#9;&#9;return Image{}, fmt.Errorf("something went wrong")
&#9;}

&#9;return r.Image, nil
}</code></pre><p>The client extends the error handling approach started by <code>preProcessError</code> by verifying whether the message included in the gRPC status matches any known error. If this lines up, a proper Go error will be returned, which can then be checked by downstream consumers.</p><div><hr></div><p>Restoring errors is an easy method to hide implementation details in scenarios where requests are sent over a network. This way, consumers can interact with interfaces as intended. Improved support for narrowing error types in future Go releases would be beneficial for building robust error handling. Until then, documentation is the best we can do to ensure our code behaves as expected.</p>]]></content:encoded></item><item><title><![CDATA[Harnessing the Power of Go Interfaces for Decoupling and Scaling at Inngest]]></title><description><![CDATA[Interfaces hide implementation details by specifying the capabilities or behavior of an object.]]></description><link>https://www.brunoscheufler.com/p/2024-06-29-harnessing-the-power-of-go-interfaces</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-06-29-harnessing-the-power-of-go-interfaces</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sat, 29 Jun 2024 11:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Interfaces hide implementation details by specifying the capabilities or behavior of an object. This concept has appeared in different programming languages over time, using different terms like <em><a href="https://doc.rust-lang.org/book/ch10-02-traits.html?ref=brunoscheufler.com">traits</a></em> or <em><a href="https://docs.swift.org/swift-book/documentation/the-swift-programming-language/protocols/?ref=brunoscheufler.com">protocols</a></em>.</p><p>The same rules apply in Go, take the classic example of <a href="https://go.dev/tour/methods/21?ref=brunoscheufler.com"><code>io.Reader</code></a>:</p><pre><code>type Reader interface {
&#9;Read(p []byte) (n int, err error)
}</code></pre><p>This interface is used <em>everywhere</em>. It doesn&#8217;t matter if you&#8217;re reading from a file, a network socket, or a data structure in the same process: As long as you need to read bytes from a source, the interface can be used.</p><p>This is amazing because it allows decoupling the implementation from the downstream consumers. Your code shouldn&#8217;t care about <em>how</em> exactly a file is read from disk, or whether it is read from disk at all.</p><p>Interfaces enable many use cases, and in Go, they&#8217;re very idiomatic to use whenever it makes sense. One rather interesting case I&#8217;ve run into recently, was decoupling data access in our production codebase at Inngest.</p><h2>An intro to Inngest</h2><p>Inngest orchestrates durable functions and workflows, hosted on any service that can accept HTTP requests. Each function includes one or more steps triggered by events.</p><p>During a function run, we need to store details like the return value of your steps so we can load them on the next step. Here&#8217;s an example function that showcases Inngest&#8217;s capabilities: When a user signs up, Inngest runs steps to load the user data, send a welcome email, and wait for a post creation event, triggering further actions based on the outcome.</p><pre><code>export default inngest.createFunction(
  { id: "activation-email" },
  { event: "app/user.created" },
  async ({ event, step }) =&gt; {
&#9;  const user = await step.run("load-user", async () =&gt; {
      return await db.loadUser({ email: event.user.email });
    });

    await step.run("send-welcome-email", async () =&gt; {
      return await sendWelcomeEmail({ email: user.email, name: user.name });
    });

    // Wait for an "app/post.created" event
    const postCreated = await step.waitForEvent("wait-for-post-creation", {
      event: "app/post.created",
      match: "data.user.id", // the field "data.user.id" must match
      timeout: "24h", // wait at most 24 hours
    });

    if (!postCreated) {
      // If no post was created, send a reminder email
      await step.run("send-reminder-email", async () =&gt; {
        return await sendReminderEmail({
          email: event.user.email,
          name: user.name,
        });
      });
    }
  }
);</code></pre><h2>Storing Run State</h2><p>To handle long-running functions, Inngest stores state details like loaded user data. With waitForEvent potentially delaying execution for extended periods, our system must manage state efficiently.</p><p>Previously, our execution logic was tightly coupled to Redis for storing function run state. As Inngest Cloud grew, we needed scalable solutions like sharding and caching. We <a href="https://github.com/inngest/inngest/pull/1325?ref=brunoscheufler.com">unified</a> state access behind a <a href="https://github.com/inngest/inngest/blob/ebe92014e2548714956a058913c91b996022b1d3/pkg/execution/state/v2/interfaces.go?ref=brunoscheufler.com"><code>state.RunService</code></a> interface, ensuring flexibility and scalability.</p><h2>Consolidating Data Store Access</h2><p>By creating new interfaces for state management, we created strong guarantees about the expected functionality that each implementation needed to follow. After merging this change, we started with the real work to prepare our system for the next stage.</p><p>To gracefully handle increasing system load, we needed to address two primary concerns:</p><ul><li><p>Too many connections to the backing data store</p></li><li><p>Running out of storage capacity</p></li></ul><p>High connection counts can strain systems like our highly-available data store used to persist function run state in our Cloud environment. This is a known problem. For example, <a href="https://www.postgresql.org/docs/current/connect-estab.html?ref=brunoscheufler.com">Postgres</a> creates a new backend process every time a connection is requested.</p><p>The solution here is to use a connection pooler like PgBouncer which acts as a proxy between your application and the Postgres instance. PgBouncer simply maintains a pool of opened connections, ready to be used for incoming client connections. Once all service connections are in use, new client connections need to wait until one server connection completes.</p><p>You might think &#8220;Why don&#8217;t you just use an application-level pooler?&#8221;, and that&#8217;s a great question! Instead of running one PgBouncer instance, you might initialize a connection pool in each application process, preventing it from establishing an unbounded number of connections. This works great as long as you limit the number of application processes. Once you create new services or auto-scale your production deployment, you might establish more connections than you&#8217;ve ever planned for. That&#8217;s why this approach is very limited in large, distributed systems.</p><p>To manage increasing connections to our data store and prevent storage issues, we introduced a coordinating service. This service, implementing the state interface, allowed us to handle connections efficiently and scale seamlessly.</p><p>Previously, replacing all state access logic with a new implementation would have taken weeks, not to mention the risk we would have taken in potentially missing one or two locations in our codebase.</p><p>With the new interface, we simply swapped the direct connection logic for the coordinator service in our production environment. Admittedly, we added more precautions for rolling out to all customers, but the idea was the same.</p><p>Interfaces saved the day.</p><div><hr></div><p>Thanks a ton to Darwin, Jack, and Tony at Inngest for supporting me on this effort, it&#8217;s been super exciting to future-proof one of the most important pieces of our infrastructure.</p><p>If you&#8217;re looking for ways to build internal tools and workflows or want to implement multi-tenant queuing, please give <a href="https://www.inngest.com/?ref=brunoscheufler.com">Inngest</a> a try!</p>]]></content:encoded></item><item><title><![CDATA[Zero-Downtime Migrations in Producer/Consumer Systems]]></title><description><![CDATA[Imagine you&#8217;re building an app and want to ensure that requests don&#8217;t feel slow for the end user.]]></description><link>https://www.brunoscheufler.com/p/2024-05-19-zero-downtime-migrations-in-producer-consumer-systems</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-05-19-zero-downtime-migrations-in-producer-consumer-systems</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 19 May 2024 11:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Imagine you&#8217;re building an app and want to ensure that requests don&#8217;t feel slow for the end user. Let&#8217;s say you&#8217;re working on a signup flow that needs to send an activation email and update external systems like billing and a CRM. You don&#8217;t really <em>need</em> to block the request until all these steps have run to completion. Instead, you could decouple this workflow into a producer and consumer. When receiving the signup request, your API would simply enqueue a task to some distributed queue, which a worker service could then pick up and process.</p><p>After a couple of months, you&#8217;ve shipped the initial version of this workflow and requirements have changed. Unfortunately, the worker processing the messages needs to be updated to handle a slightly different format.</p><p>You&#8217;ve thought long and hard and decided to prepare the following rollout plan: For a short transition period, you&#8217;ll ensure that the system keeps processing both old and new messages. This is done by updating the consumer to handle both the new and old message formats. The producer is updated to provide the new format. After all old messages have been processed, you can remove the code handling the old format.</p><p>You prepare a pull request including updates to both the producer and consumer and all tests pass. You feel good about handling the problem and deploy to production. You&#8217;re using a managed container service and watch while new instances are spinning up. Suddenly, your error rate starts spiking. Messages are failing in the consumer. You begin to sweat a little and start exploring the current logs to find consumer containers complaining about malformed messages. Of course, that&#8217;s why you added graceful handling in the pull request! You assume that the new code is broken and decide to roll back. For some reason, the errors keep coming in.</p><p>What&#8217;s going on? Let&#8217;s rewind.</p><p>Orchestration systems like Kubernetes usually deploy new containers side by side with old instances. This allows for a graceful, zero-downtime rollout. This also means that two versions of your code are running side by side.</p><p>As distributed systems are notoriously hard to predict, there&#8217;s no way that you can or should be timing the deployment of different components in one deployment. That&#8217;s why you can&#8217;t be certain that all consumers are deployed before producers start rolling out.</p><p>Let&#8217;s go back to the time of deploying your changes.</p><p>After sifting through the errors some more, you notice that all logs originate from <em>old</em> containers. Suddenly, everything starts to make sense!</p><p>You roll forward to the latest version, wait for a couple of minutes until all old pods have terminated, and the errors disappear.</p><p>Let&#8217;s recap what went wrong here. First, you expected your changes to roll out in a certain order. With most systems, you do not get these guarantees. Second, you did not handle the case where old consumers were picking up messages enqueued by new producers. You have to be <em>really really</em> careful that new consumers handle old messages <em>and</em> old consumers handle new messages. If either invariant is violated, you&#8217;re potentially breaking your system.</p><p>So how could this have been avoided?</p><p>One possible solution could have been to separate the deployment of producers and consumers. You could have split up your pull request to isolate the changes enabling graceful handling of old messages to the consumer. After deploying this and waiting for all old consumers to terminate, you could have followed up with another deployment updating the producer code to publish new messages. At that point in time, no new messages could have physically reached old consumers.</p><p>In any case, the system should have gracefully retried the failing messages. After a couple of attempts, a message should have reached a new consumer and succeeded, even if the initial attempt failed. As long as you never drop a message that wasn&#8217;t processed yet, the system is able to recover (albeit after some downtime).</p><p>Rolling out changes in distributed systems is a tricky problem because your tests won&#8217;t catch this kind of inconsistency. Tests run on the current version of the codebase, which compiles just fine. <em>Ideally</em>, your test runner would know about old code running in production and test producer changes against old and new builds.</p><div><hr></div><p>Instead of writing all of this yourself, you should be using a solution like <a href="https://www.inngest.com/?ref=brunoscheufler.com">Inngest</a>, which happens to be my current employer.</p>]]></content:encoded></item><item><title><![CDATA[Derisking Product Rollouts with Feature Flags]]></title><description><![CDATA[In most successful companies, products and their underlying infrastructure are constantly evolving to changing needs and growing demand.]]></description><link>https://www.brunoscheufler.com/p/2024-05-19-derisking-product-rollouts-with-feature-flags</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-05-19-derisking-product-rollouts-with-feature-flags</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 19 May 2024 10:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In most successful companies, products and their underlying infrastructure are constantly evolving to changing needs and growing demand.</p><p>Shipping new features, improvements, or other changes can require careful planning ahead of a release. While you can try to imagine and rule out different failure scenarios ahead of time, you never really know what&#8217;s going to happen until you ship.</p><p>Sure, you could set up shadow reads/writes or split production traffic to hit an isolated clone running the new code, but depending on your resource budget and current infrastructure, that effort might be hard to justify.</p><p>For new versions of an existing feature, you could also fall back to the old logic on failure or compare new output with existing results and log any deviations, as they occur.</p><p>In small, fast-moving teams, you want to roll out features at a faster pace, but you cannot afford to break the experience for higher-paying and enterprise customers. So what if you could gradually roll out new features across segments, starting with free users and moving up the subscription tiers in steps?</p><p>Feature flags are a tool used for toggling specific parts of your product on or off for a certain set of users without having to reboot or reconfigure services. With feature flags, you can distinguish between different segments (groups of users) and gradually enable a flag for a small and then growing percentage of users over time. If something goes wrong, you can simply disable the flag and users will work on the previous version.</p><p>This flow only works when you can switch between implementations without migrations. In case you&#8217;re releasing a big feature by migrating individual users, consider enrolling eligible accounts in the migration process incrementally. You should probably strive to limit large migrations to speed up the rollout process and reduce your team&#8217;s maintenance burden.</p><p>In case you&#8217;re switching between two implementations with identical interfaces, feature flags are a perfect match. Slowly increase the percentage of users exposed to the new implementation and optionally enable fallbacks to the previous version for all users. Once everyone is using the new code, you can slowly remove the fallbacks as long as the new implementation doesn&#8217;t cause any issues. Once nobody is using the fallback, you can completely drop the old code.</p><p>In the past, I&#8217;ve used LaunchDarkly to create and integrate feature flags into products I worked on. With LaunchDarkly, you can configure feature flags to target specific users in different segments (or a percentage thereof), without worrying about performance degradation, as flag rules are stored in memory and evaluated without a network roundtrip.</p><div><hr></div><p>Properly planning feature rollouts removes a lot of stress factors from the team. If you have a clear plan and options to roll back to the previous experience if anything goes wrong, your coworkers and customers will thank you for it.</p>]]></content:encoded></item><item><title><![CDATA[Building a DNS message parser in Go]]></title><description><![CDATA[The internet has grown from a couple of research lab sites to a vibrant ecosystem of content and services.]]></description><link>https://www.brunoscheufler.com/p/2024-05-12-building-a-dns-message-parser</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-05-12-building-a-dns-message-parser</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 12 May 2024 10:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The internet has grown from a couple of research lab sites to a vibrant ecosystem of content and services. Without the ability to resolve human-readable names to machine identifiers, however, it&#8217;d be impossible to find anything in the ever-growing web of interconnected systems.</p><p>In November 1987, <a href="https://datatracker.ietf.org/doc/html/rfc1034?ref=brunoscheufler.com">RFC 1034</a> and <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com">RFC 1035</a> first defined the domain name system (DNS), a distributed database containing the names of resources, designed to be extensible beyond the initial scope. Software including browsers and other internet-enabled apps refers to a DNS resolver, which queries name servers to resolve hostnames to IP addresses, mail settings, or other information.</p><p>To facilitate communication between DNS resolvers and name servers, RFC 1035 specifies the DNS wire format. DNS messages are transmitted as a series of octets (bytes). In this guide, we&#8217;ll try to understand what goes into a typical DNS query and response by implementing a DNS message parser in Go.</p><pre><code> 0               1
 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|       1       |       2       |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|       3       |       4       |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|       5       |       6       |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+</code></pre><p>Before we get to parsing messages, let&#8217;s look at some fundamental concepts used through DNS, including domain names and resource records.</p><h2>Domain Names</h2><pre><code>&lt;domain&gt; ::= &lt;subdomain&gt; | " "

&lt;subdomain&gt; ::= &lt;label&gt; | &lt;subdomain&gt; "." &lt;label&gt;

&lt;label&gt; ::= &lt;letter&gt; [ [ &lt;ldh-str&gt; ] &lt;let-dig&gt; ]

&lt;ldh-str&gt; ::= &lt;let-dig-hyp&gt; | &lt;let-dig-hyp&gt; &lt;ldh-str&gt;

&lt;let-dig-hyp&gt; ::= &lt;let-dig&gt; | "-"

&lt;let-dig&gt; ::= &lt;letter&gt; | &lt;digit&gt;

&lt;letter&gt; ::= any one of the 52 alphabetic characters A through Z in
upper case and a through z in lower case

&lt;digit&gt; ::= any one of the ten digits 0 through 9</code></pre><p>Domain Names as defined in <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-2.3.1">Section 2.3.1</a> and <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-3.1">Section 3.1.</a> are represented as sequences of labels separated by dots. Labels must start with letters followed by optional numbers and can contain, but not end with hyphens. Each label is represented as a one-octet length field followed by that number of octets.</p><p>Every domain name ends with the null label of the root (visualized using a trailing dot), which is why a domain name is terminated by a zero-length octet.</p><p>To accommodate for <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-4.1.4">message compression</a>, the first two high-order bits of every label&#8217;s length octet must be zero. In case compression is used, both bits are 1, followed by an offset value pointing to another label or domain name.</p><pre><code>func parseDomainName(fullMessage []byte, bytes []byte) (string, []byte) {
&#9;domainName := ""
&#9;for {
&#9;&#9;length := bytes[0]

&#9;&#9;isCompressed := length&amp;0b1100_0000 &gt; 0
&#9;&#9;if isCompressed {
&#9;&#9;&#9;pointer := (bytes[0] &lt;&lt; 2) + bytes[1]

&#9;&#9;&#9;resolvedPointer, _ := parseDomainName(fullMessage, fullMessage[pointer:])
&#9;&#9;&#9;domainName += resolvedPointer

&#9;&#9;&#9;bytes = bytes[2:]

&#9;&#9;&#9;break
&#9;&#9;}

&#9;&#9;bytes = bytes[1:]

&#9;&#9;if length == 0 {
&#9;&#9;&#9;break
&#9;&#9;}

&#9;&#9;label := bytes[0:length]

&#9;&#9;domainName += string(label)
&#9;&#9;domainName += "."

&#9;&#9;bytes = bytes[length:]
&#9;}

&#9;return domainName, bytes
}</code></pre><h2>Resource Records (RRs)</h2><p>As I&#8217;ve said before, the domain name system was created as a database to translate resource names to information like IP addresses and other values. The atomic unit in this database is referred to as a Resource Record (RR) and is defined in Section <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-3.2.1">3.2.1.</a> Each RR is addressed using a name (as defined above) and stores resource data (RDATA) for a given resource type (TYPE) and class (CLASS), expiring after the TTL.</p><pre><code>                                    1  1  1  1  1  1
      0  1  2  3  4  5  6  7  8  9  0  1  2  3  4  5
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                                               |
    /                                               /
    /                      NAME                     /
    |                                               |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                      TYPE                     |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                     CLASS                     |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                      TTL                      |
    |                                               |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                   RDLENGTH                    |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--|
    /                     RDATA                     /
    /                                               /
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+</code></pre><p>Resource Records exist for different data types like IPv4 (<a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-3.4.1">A records</a>) and IPv6 (defined in <a href="https://datatracker.ietf.org/doc/html/rfc3596?ref=brunoscheufler.com">RFC 3596</a>) addresses, as well as text (<a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-3.3.14">TXT</a>) and mail-related (<a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-3.3.9">MX</a>) records.</p><pre><code>type RDataCNAME struct {
&#9;CName string // &lt;domain-name&gt; which specifies canonical or primary name for owner, owner name is alias
}

func (d RDataCNAME) String() string {
&#9;return d.CName
}

type RDataMX struct {
&#9;Preference uint16 // lower values are preferred
&#9;Exchange   string // &lt;domain-name&gt; of host willing to act as mail exchange
}

func (d RDataMX) String() string {
&#9;return fmt.Sprintf("%d %s", d.Preference, d.Exchange)
}

type RDataNS struct {
&#9;NsdName string // &lt;domain-name&gt; specifying host which should be authoritative for specified class
}

func (d RDataNS) String() string {
&#9;return d.NsdName
}

type RDataTXT struct {
&#9;TxtData string // one or more character strings
}

func (d RDataTXT) String() string {
&#9;return d.TxtData
}

type RDataA struct {
&#9;Address string // host internet address
}

func (d RDataA) String() string {
&#9;return d.Address
}

type RDataAAAA struct { // RFC 3596
&#9;Address string // 128-bit IPv6 address
}

func (d RDataAAAA) String() string {
&#9;return d.Address
}

type ResourceRecord struct {
&#9;Name     string       // domain name associated with record
&#9;Type     RecordType   // two octets containing RR TYPE code
&#9;Class    RecordClass  // two octets containing RR class code
&#9;TTL      uint32       // time in seconds until record should be refreshed in cache
&#9;RDLength uint16       // length of octets in RData
&#9;RData    fmt.Stringer // variable-length string of octets describing resource, format varies by TYPE and CLASS
}</code></pre><pre><code>// transparently handles empty sections
func parseResourceRecords(fullMessage []byte, messageBytes []byte, numRecords uint16) ([]ResourceRecord, []byte, error) {
&#9;resourceRecords := make([]ResourceRecord, numRecords)

&#9;for i := range resourceRecords {

&#9;&#9;var domainName string
&#9;&#9;domainName, messageBytes = parseDomainName(fullMessage, messageBytes)

&#9;&#9;resourceRecords[i].Name = domainName

&#9;&#9;resourceRecords[i].Type = RecordType(binary.BigEndian.Uint16(messageBytes[0:2]))
&#9;&#9;messageBytes = messageBytes[2:]

&#9;&#9;resourceRecords[i].Class = RecordClass(binary.BigEndian.Uint16(messageBytes[0:2]))
&#9;&#9;messageBytes = messageBytes[2:]

&#9;&#9;resourceRecords[i].TTL = binary.BigEndian.Uint32(messageBytes[0:4])
&#9;&#9;messageBytes = messageBytes[4:]

&#9;&#9;resourceRecords[i].RDLength = binary.BigEndian.Uint16(messageBytes[0:2])
&#9;&#9;messageBytes = messageBytes[2:]

&#9;&#9;switch resourceRecords[i].Type {
&#9;&#9;case TypeA:
&#9;&#9;&#9;// read first 32bit/4 byte
&#9;&#9;&#9;ipv4 := net.IPAddr{
&#9;&#9;&#9;&#9;IP: messageBytes[0:4],
&#9;&#9;&#9;}
&#9;&#9;&#9;messageBytes = messageBytes[4:]

&#9;&#9;&#9;resourceRecords[i].RData = RDataA{
&#9;&#9;&#9;&#9;Address: ipv4.String(),
&#9;&#9;&#9;}
&#9;&#9;case TypeAAAA:
&#9;&#9;&#9;// read first 128bit (https://datatracker.ietf.org/doc/html/rfc3596#section-2.2)
&#9;&#9;&#9;ipv6 := net.IPAddr{
&#9;&#9;&#9;&#9;IP: messageBytes[0:16],
&#9;&#9;&#9;}
&#9;&#9;&#9;messageBytes = messageBytes[16:]

&#9;&#9;&#9;resourceRecords[i].RData = RDataAAAA{
&#9;&#9;&#9;&#9;Address: ipv6.String(),
&#9;&#9;&#9;}
&#9;&#9;case TypeCNAME:
&#9;&#9;&#9;var canonicalDomainName string
&#9;&#9;&#9;canonicalDomainName, messageBytes = parseDomainName(fullMessage, messageBytes)

&#9;&#9;&#9;resourceRecords[i].RData = RDataCNAME{
&#9;&#9;&#9;&#9;CName: canonicalDomainName,
&#9;&#9;&#9;}
&#9;&#9;case TypeMX:
&#9;&#9;&#9;rDataStr := string(messageBytes[0:resourceRecords[i].RDLength])
&#9;&#9;&#9;messageBytes = messageBytes[resourceRecords[i].RDLength:]

&#9;&#9;&#9;preference := binary.BigEndian.Uint16([]byte(rDataStr)[0:2])
&#9;&#9;&#9;var exchangeDomainName string
&#9;&#9;&#9;exchangeDomainName, messageBytes = parseDomainName(fullMessage, []byte(rDataStr[2:]))

&#9;&#9;&#9;resourceRecords[i].RData = RDataMX{
&#9;&#9;&#9;&#9;Preference: preference,
&#9;&#9;&#9;&#9;Exchange:   exchangeDomainName,
&#9;&#9;&#9;}
&#9;&#9;case TypeNS:
&#9;&#9;&#9;rDataStr := string(messageBytes[0:resourceRecords[i].RDLength])
&#9;&#9;&#9;messageBytes = messageBytes[resourceRecords[i].RDLength:]

&#9;&#9;&#9;var nsDomainName string
&#9;&#9;&#9;nsDomainName, messageBytes = parseDomainName(fullMessage, []byte(rDataStr))

&#9;&#9;&#9;resourceRecords[i].RData = RDataNS{
&#9;&#9;&#9;&#9;NsdName: nsDomainName,
&#9;&#9;&#9;}
&#9;&#9;case TypeTXT:
&#9;&#9;&#9;rDataStr := string(messageBytes[0:resourceRecords[i].RDLength])
&#9;&#9;&#9;messageBytes = messageBytes[resourceRecords[i].RDLength:]

&#9;&#9;&#9;resourceRecords[i].RData = RDataTXT{
&#9;&#9;&#9;&#9;TxtData: rDataStr,
&#9;&#9;&#9;}
&#9;&#9;}
&#9;}

&#9;return resourceRecords, messageBytes, nil
}</code></pre><p>On its own, resource records are not very useful: We want clients to be able to query DNS name servers to retrieve relevant records by specifying a domain name, record class, and type. To transmit queries like this, the RFC introduces a common format for messages.</p><h2>Messages</h2><pre><code>    +---------------------+
    |        Header       |
    +---------------------+
    |       Question      | the question for the name server
    +---------------------+
    |        Answer       | RRs answering the question
    +---------------------+
    |      Authority      | RRs pointing toward an authority
    +---------------------+
    |      Additional     | RRs holding additional information
    +---------------------+</code></pre><p>Both queries and responses are transmitted using messages, as defined in <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-4.1">Section 4.1</a>. Messages start with a header, followed by one or more questions, resource records for answers, pointing to authoritative name servers (authority), and related information (additional) which are not strictly answers to the question.</p><pre><code>type Message struct {
&#9;Header     MessageHeader    // always present
&#9;Question   []QuestionEntry  // question for name server
&#9;Answer     []ResourceRecord // RRs answering question
&#9;Authority  []ResourceRecord // RRs pointing towards authority
&#9;Additional []ResourceRecord // RRs holding additional info
}</code></pre><p>To parse the entire message, we must start with the header and work our way through the bytes.</p><pre><code>func parseMessage(messageBytes []byte) (Message, error) {
&#9;fullMessage := messageBytes

&#9;headerBytes := messageBytes[0:headerSizeBytes]
&#9;header, err := parseMessageHeader(headerBytes)
&#9;if err != nil {
&#9;&#9;return Message{}, fmt.Errorf("could not parse header: %w", err)
&#9;}

&#9;messageBytes = messageBytes[headerSizeBytes:]
&#9;questionEntries, messageBytes, err := parseQuestions(fullMessage, messageBytes, header.QDCount)
&#9;if err != nil {
&#9;&#9;return Message{}, fmt.Errorf("could not parse questions: %w", err)
&#9;}

&#9;answerRRs, messageBytes, err := parseResourceRecords(fullMessage, messageBytes, header.ANCount)
&#9;if err != nil {
&#9;&#9;return Message{}, fmt.Errorf("could not parse answer RRs: %w", err)
&#9;}

&#9;authorityRRs, messageBytes, err := parseResourceRecords(fullMessage, messageBytes, header.NSCount)
&#9;if err != nil {
&#9;&#9;return Message{}, fmt.Errorf("could not parse authority RRs: %w", err)
&#9;}

&#9;additionalRRs, messageBytes, err := parseResourceRecords(fullMessage, messageBytes, header.ARCount)
&#9;if err != nil {
&#9;&#9;return Message{}, fmt.Errorf("could not parse additional RRs: %w", err)
&#9;}

&#9;return Message{
&#9;&#9;Header:     header,
&#9;&#9;Question:   questionEntries,
&#9;&#9;Answer:     answerRRs,
&#9;&#9;Authority:  authorityRRs,
&#9;&#9;Additional: additionalRRs,
&#9;}, nil
}</code></pre><p>Let&#8217;s start by parsing the message header!</p><h3>Header</h3><pre><code>
                                    1  1  1  1  1  1
      0  1  2  3  4  5  6  7  8  9  0  1  2  3  4  5
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                      ID                       |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |QR|   Opcode  |AA|TC|RD|RA|   Z    |   RCODE   |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                    QDCOUNT                    |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                    ANCOUNT                    |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                    NSCOUNT                    |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                    ARCOUNT                    |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+</code></pre><p>The header section is defined in Section <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-4.1.1">4.1.1.</a> of RFC 1035 and includes important information to identify and match the current response to a client request, determine whether the current message is a request or response (QR), specify a query kind (OPCODE), response code (RCODE), and other important flags.</p><pre><code>type MessageHeader struct {
&#9;ID      uint16   // identifier assigned by program that generated query, copied to reply
&#9;QR      bool     // one-bit field that specifies if message is query (0, false) or response (1, true)
&#9;OPCode  OpCode   // kind of query
&#9;AA      bool     // only in response, specifies that responding server is authority for domain name (corresponds to name matching query or first owner name in answer)
&#9;TC      bool     // TrunCation, specifies message was truncated due to length greater than permitted on the transmission channel
&#9;RD      bool     // Recursion Desired: set in query, copied to response, directs name server to pursue query recursively
&#9;RA      bool     // Recursion Available, set (1) or cleared (0) in response, denotes whether name server supports recursive queries
&#9;Z       struct{} // reserved, must be 0
&#9;RCode   RCode    // response code
&#9;QDCount uint16   // number of entries in questions section
&#9;ANCount uint16   // number of RRs in answer section
&#9;NSCount uint16   // number of RRs in authority records section
&#9;ARCount uint16   // number of RRs in additional records section
}</code></pre><p>From the specification, we already know the header to be 2 octets * 6 rows = 96 bits = 12 bytes in length. If this doesn&#8217;t match up, we can return an error.</p><p>With some bitwise operators, we can extract the relevant values from our structure, shifting bits whenever we need to access misaligned values in the octets. Using the <a href="https://pkg.go.dev/encoding/binary?ref=brunoscheufler.com">encoding/binary</a> package, we can translate byte sequences to numbers, which is used for parsing the ID and count values.</p><pre><code>const headerSizeBytes = (8 * 2 * 6) / 8 // two octets * 6

func parseMessageHeader(headerBytes []byte) (MessageHeader, error) {
&#9;if len(headerBytes) != 12 {
&#9;&#9;return MessageHeader{}, fmt.Errorf("expected header to be 12 bytes")
&#9;}

&#9;id := binary.BigEndian.Uint16(headerBytes[0:2]) // first two bytes = 16 bits = 2 octets

&#9;secondRow := headerBytes[2:4] // second two bytes = 16 bits = 2 octets

&#9;qr := secondRow[0]&amp;byte(0b1000_0000) &gt; 0

&#9;// after consuming QR, shift entire octet to the left so that opcode takes up leading bits
&#9;secondRow[0] = secondRow[0] &lt;&lt; 1

&#9;// retrieve four-bit opcode and shift right by remaining 4 bits in octet to "index" at 0 instead of 8
&#9;opcode := OpCode(secondRow[0] &amp; byte(0b1111_0000) &gt;&gt; 4)

&#9;aa := secondRow[0]&amp;byte(0b0000_1000) &gt; 0
&#9;tc := secondRow[0]&amp;byte(0b0000_0100) &gt; 0
&#9;rd := secondRow[0]&amp;byte(0b0000_0010) &gt; 0

&#9;ra := secondRow[1]&amp;byte(0b1000_0000) &gt; 0

&#9;rcode := RCode(secondRow[1] &amp; byte(0b0000_1111))

&#9;qdcount := binary.BigEndian.Uint16(headerBytes[4:6])
&#9;ancount := binary.BigEndian.Uint16(headerBytes[6:8])
&#9;nscount := binary.BigEndian.Uint16(headerBytes[8:10])
&#9;arcount := binary.BigEndian.Uint16(headerBytes[10:12])

&#9;return MessageHeader{
&#9;&#9;ID:      id,
&#9;&#9;QR:      qr,
&#9;&#9;OPCode:  opcode,
&#9;&#9;AA:      aa,
&#9;&#9;TC:      tc,
&#9;&#9;RD:      rd,
&#9;&#9;RA:      ra,
&#9;&#9;Z:       struct{}{},
&#9;&#9;RCode:   rcode,
&#9;&#9;QDCount: qdcount,
&#9;&#9;ANCount: ancount,
&#9;&#9;NSCount: nscount,
&#9;&#9;ARCount: arcount,
&#9;}, nil
}</code></pre><p>Once we have parsed the message header, we know how many questions, and resource records we can expect.</p><h3>Question</h3><pre><code>                                    1  1  1  1  1  1
      0  1  2  3  4  5  6  7  8  9  0  1  2  3  4  5
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                                               |
    /                     QNAME                     /
    /                                               /
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                     QTYPE                     |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+
    |                     QCLASS                    |
    +--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+</code></pre><p>Questions are defined in <a href="https://datatracker.ietf.org/doc/html/rfc1035?ref=brunoscheufler.com#section-4.1.2">Section 4.1.2.</a> and consist of a domain name (QNAME), a query type (QTYPE), and class (QCLASS).</p><pre><code>func parseQuestions(fullMessage []byte, messageBytes []byte, numQuestions uint16) ([]QuestionEntry, []byte, error) {
&#9;questionEntries := make([]QuestionEntry, numQuestions)

&#9;for i := range questionEntries {
&#9;&#9;var domainName string
&#9;&#9;domainName, messageBytes = parseDomainName(fullMessage, messageBytes)
&#9;&#9;questionEntries[i].QName = domainName

&#9;&#9;questionEntries[i].QType = QType(binary.BigEndian.Uint16(messageBytes[0:2]))
&#9;&#9;messageBytes = messageBytes[2:]

&#9;&#9;questionEntries[i].QClass = QClass(binary.BigEndian.Uint16(messageBytes[0:2]))
&#9;&#9;messageBytes = messageBytes[2:]
&#9;}

&#9;return questionEntries, messageBytes, nil
}</code></pre><div><hr></div><p>This is all it takes to build a DNS message parser! Admittedly, it doesn&#8217;t handle cool extensions like DNSSEC, but you could go ahead and extend it easily. This goes to show that the original idea of an extensible naming system has lived up to its promises, fulfilling a critical role in powering internet infrastructure for decades to come.</p><p>I was pleasantly surprised how straightforward it was to follow the RFC and implement the parser in Go. If I find the time, I might look at other RFCs and try to implement more protocols. If you have any suggestions or questions, feel free to <a href="mailto:bruno@brunoscheufler.com">send a mail</a>!</p>]]></content:encoded></item><item><title><![CDATA[Joining Inngest]]></title><description><![CDATA[In the coming decades, there&#8217;s only going to be more software to build, operate, and scale for platforms we&#8217;re only just inventing, with code running on nearly everything we interact with day to day.]]></description><link>https://www.brunoscheufler.com/p/2024-04-23-joining-inngest</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-04-23-joining-inngest</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Tue, 23 Apr 2024 10:00:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/be0a2d99-4919-42c4-a75b-9dc27f0b0191_2400x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ow_i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ow_i!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 424w, https://substackcdn.com/image/fetch/$s_!ow_i!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 848w, https://substackcdn.com/image/fetch/$s_!ow_i!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 1272w, https://substackcdn.com/image/fetch/$s_!ow_i!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ow_i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Joining Inngest&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Joining Inngest" title="Joining Inngest" srcset="https://substackcdn.com/image/fetch/$s_!ow_i!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 424w, https://substackcdn.com/image/fetch/$s_!ow_i!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 848w, https://substackcdn.com/image/fetch/$s_!ow_i!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 1272w, https://substackcdn.com/image/fetch/$s_!ow_i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3025c3f-988c-4911-98cc-3d4f5029aa6a_2400x800.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><p>In the coming decades, there&#8217;s only going to be more software to build, operate, and scale for platforms we&#8217;re only just inventing, with code running on nearly everything we interact with day to day.</p><p>To help advance the current frontier into AI, which is built on massively distributed computing, understanding the engineering and technology stack from end to end has never been as important as today.</p><p>I&#8217;ve always had a deep interest in understanding how businesses and technology work at the foundational level. That&#8217;s why I joined the team at <a href="https://hygraph.com/?ref=brunoscheufler.com">Hygraph</a> when they were 5 people building out of a small office. I took the opportunity to learn the software engineering stack for massively multi-tenant B2B SaaS products. I helped build globally-distributed and highly-available cloud infrastructure, and I got to understand what it takes to scale an engineering organization from Pre-Seed through Series B, taking care of hiring and onboarding processes, documentation, and developer experience among other areas.</p><p>At TUM, I got insights into computer science and business fundamentals, getting into operating systems and computer architecture, distributed systems and <a href="https://dse.in.tum.de/?ref=brunoscheufler.com">heterogeneous computing</a> with FPGAs, as well as corporate finance, business law, and operations research. I found it invaluable to discover areas that I wouldn&#8217;t have explored on my own.</p><p>After completing my studies and having built a couple of side projects over the years, I ramped up efforts in starting my own company together with my long-time friend, <a href="https://timweiss.net/?ref=brunoscheufler.com">Tim</a>. We believed the timing was spot on, and we wanted to solve problems we experienced ourselves in past jobs. First, we attempted to validate Anzu, which provided essential building blocks for engineering teams wanting to move fast, followed up by CodeTrail for providing better onboarding documentation.</p><p>We grew and got better every iteration, especially in sales and marketing. We crafted well-designed products, created strongly-personalized sales sequences, and ran cold-calling sessions to get our first customers.</p><p>Unfortunately, the cards were stacked against us. We failed to identify and validate an urgent, unmet business need. We didn&#8217;t have the network to collaborate with pilot customers before writing a line of code. And we didn&#8217;t have the capital for continued exploration into other ideas.</p><p>So we quit.</p><p>After a long and candid discussion with Tim, we decided it made sense for us to go back to square one and grow both our skills and network, our credentials, and financial independence with the eventual goal of getting back to building a business. The next time we work together as founders, however, we&#8217;ll solve a real problem with real customers and the capital to sustain the grind until we reach exit velocity.</p><p>Thus, I accelerated the pace of interviewing and talking to different teams in March, seeing where I could best bring in my experience and have an outsized impact and growth opportunities. Through this process, I learned a ton already, but that&#8217;s for another time.</p><p>Today, I&#8217;m excited to announce that I&#8217;m joining <a href="https://www.inngest.com/?ref=bruno-post">Inngest</a> as a Senior Distributed Systems Engineer starting next Monday.</p><p>In the past, I&#8217;ve spent more time than I&#8217;d like to admit on decoupling critical services, managing distributed queues, handling failing jobs gracefully, and wrapping my head around distributed transactions. Inngest makes it ridiculously easy to build durable functions and background jobs, without managing any of the infrastructure. It&#8217;s almost magical.</p><p>Earlier this year, the team raised $6.1M in new funding by a16z (Andreessen Horowitz), with follow-on investment from existing investors: GGV, Afore Capital and Guillermo Rauch. I&#8217;m joining Inngest to help scale the backend infrastructure for the next growth stage and beyond.</p><p>Over the past weeks, I&#8217;ve had an exceptionally great interview process and incredible support from Tony and the entire team at Inngest on all process-related questions. I&#8217;m absolutely thrilled to get to know everyone on the team over the coming weeks and dive into the engineering process once again.</p><p>I&#8217;ll keep you posted &#128075;</p><p>A heartfelt thank you to everyone who gave their advice and recommendations over the course of this long-winding process and throughout the lows and highs, this wouldn&#8217;t have been possible without you.</p>]]></content:encoded></item><item><title><![CDATA[Building a Hybrid Search Experience]]></title><description><![CDATA[You might have noticed a recent addition to the top right corner of this blog: I&#8217;ve added a global search experience.]]></description><link>https://www.brunoscheufler.com/p/2024-04-03-building-a-hybrid-search-experience</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-04-03-building-a-hybrid-search-experience</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Wed, 03 Apr 2024 18:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You might have noticed a recent addition to the top right corner of this blog: I&#8217;ve added a global search experience. More specifically, I&#8217;ve implemented a hybrid search solution based on OpenAI Embeddings, Supabase, and PostgreSQL full-text search. In this post, I&#8217;d like to run through the design decisions I made building this system and explain how the current iteration works end to end.</p><h2>Previously</h2><p>Long-time readers may recall previous iterations of this blog that included an Algolia-powered search experience. When I switched out Gatsby in favor of Next.js for the foundation of the current stack <a href="__GHOST_URL__/blog/2020-10-20-rebuilding-my-portfolio-using-nextjs-tailwind">in 2020</a>, I removed all the bells and whistles I didn&#8217;t want to maintain anymore, including the in-house newsletter and search solutions. I wanted to focus on the essence of the blog, the content.</p><p>Over time, discoverability has become a challenge: New readers have a hard time finding content they&#8217;re curious about, and I&#8217;m losing track of what I wrote about years ago. Introducing a new search experience is one of the first steps I&#8217;m taking to improve content discoverability, with more changes to follow in the coming weeks.</p><h2>Getting to Hybrid Search</h2><p>When I started exploring the possible ways to implement search, I decided not to run with a managed solution like Algolia. Instead, I want to build a search stack tailored to the blog&#8217;s needs.</p><p>When users search for content, they might remember some detail of a post they read, or they might be interested in a certain technology. They might have come across one of my startups or projects. They might enter an area of work like system architecture, continuous integration, or containers. And they may paraphrase, use synonyms, or other related words.</p><p>In short, they will use a thousand possible ways to reach the same goal, finding a post. And I need to supply the right tools to make this task easier.</p><p>Keyword or full-text search allows finding exact matches, for instance, returning all posts mentioning CodeTrail, my most recent startup. This breaks down when you diverge from the phrasing I&#8217;ve used. Replace VM with virtualization and you&#8217;re out of luck.</p><p>Semantic search operates on learned representations of concepts in vector space. Imagine having to group similar words in a room. You might put ice closer to cold, and coffee closer to hot and warm. Containers are close to virtualization and far away from surfboard. So-called embedding models are trained to group words in an n-dimensional space minimizing the space between tokens that should be semantically similar. Comparing the embedded vector representation of a query string with the embedded vector representation of a blog post&#8217;s content using cosine similarity ranks closely-matching (approaching 1) or unrelated (approaching 0) results. Intuitively, this works because minimizing the angle in each learned dimension of the vector space corresponds with close semantic similarity.</p><p>Semantic search is a powerful tool to fetch semantically similar results for a query and the same result for semantically similar queries at the same time.</p><p>Semantic search operates on a predefined vocabulary that might not know about the latest technology. Running &#8220;Apple Vision Pro&#8221; through an embedding model trained before the last WWDC will most likely not arrange it closely to &#8220;virtual reality&#8221;. This is why semantic search should be coupled with a more traditional search strategy respecting keywords.</p><p>Hybrid search approaches combine semantic search with keyword or full-text search to get the best of both worlds. Reciprocal Rank Fusion (RRF) is an algorithm that combines two ranked result lists into one unified result set. In our implementation, we will run semantic search and full-text search separately and use RRF to create a single list of results.</p><h2>Choosing the stack</h2><p>The current blog iteration is built with Next.js, and deployed on Vercel. Every post is a Markdown file, assets like images are uploaded and distributed separately. Markdown is processed using the unified ecosystem of tools including remark and rehype and finally converted to React components and rendered in the browser.</p><p>To store full-text search and embedding vectors, I decided to create a PostgreSQL instance running Supabase with the pgvector extension enabled to store vector columns and perform semantic search using cosine similarity operators between query and embedding vectors.</p><p>To convert natural language search queries to a suitable representation in vector space, I&#8217;m using the OpenAI Embeddings endpoint with the <a href="https://openai.com/blog/new-embedding-models-and-api-updates">recently launched</a> <code>text-embedding-3-small</code> model.</p><p>I&#8217;ve built an indexer service running on every push using a GitHub Actions workflow. All search requests are sent to a Vercel Edge Function endpoint. I use PostHog to understand how the search experience is adopted.</p><p>Most of the decisions are pretty standard for a small MVP-stage project and will serve for a while. I&#8217;ve deliberately planned this feature as an experiment to improve discoverability issues, and I might replace it with a better solution in the future.</p><h2>End-to-End Implementation</h2><p>To populate the posts database with new or updated post content, I&#8217;ve created a GitHub Actions workflow running an index script on push or manual dispatch.</p><p>The indexer loads all Markdown post content from disk, extracts frontmatter, and iterates over each post. It creates a full-text search vector incorporating post title, slug, keywords, and other metadata. This procedure can be changed in the future.</p><p>To prevent generating vector embeddings for unchanged posts, we calculate a hash of the current file content and compare it to the current post entry in the database. If the post hasn&#8217;t been indexed previously or hashes don&#8217;t match, we create vector embeddings for the post.</p><p>To support longer content, we split the Markdown content into chunks using langchain&#8217;s MarkdownTextSplitter. We then create embeddings for each chunk using OpenAI&#8217;s <code>text-embedding-3-small</code> model with just 512 dimensions compared to the legacy <code>text-embedding-ada-002</code>'s 1536 dimensions while achieving better performance.</p><p>Finally, we insert post embeddings into the database using the vector column data type provided by pgvector.</p><p>At runtime, when a user enters a search query, Next.js sends a client-side <code>GET</code> request to the search endpoint deployed using Vercel Edge Functions. We try to cache responses to the same query to reduce load, if possible. The function generates vector embeddings for the query string to calculate cosine similarity later on. We also generate a full-text search vector using <code>websearch_to_tsquery</code>.</p><p>Finally, we invoke the hybrid search routine, which separately finds the best matches for semantic and full-text search and applies RRF to calculate a unified result set. We also apply user filters on keywords, publishing time, etc.</p><h2>Moving to Production</h2><p>To provide a predictable experience, I&#8217;ve instrumented both frontend and backend, adding PostHog to analyze search usage for future improvements. I&#8217;ve also added service quotas to prevent adversarial users from exhausting system resources and configured billing limits in third-party services to prevent unexpected overspending.</p><h2>Iterations</h2><p>Initially, I created a minimal search experience using only semantic search. This worked well for semantically similar queries but returned suboptimal results when adding little-known entity names. Adding keyword search and combining results using RRF solved this issue.</p><p>To help users find specific content more easily, I enabled filtering support on topics, which are used throughout the blog to tag areas of work, programming languages, tools, and more.</p><p>In the future, I&#8217;ll probably add more filters and tune hybrid search parameters to perform better. This should make it even easier to sift through all posts to find the needle in the haystack.</p><div><hr></div><p>Thanks for reading all the way through this! This project has been an interesting challenge, building a usable search experience takes more effort than I initially expected. The first public release is now live, feel free to try it by clicking the search button on the top right or hitting &#8984;K.</p>]]></content:encoded></item><item><title><![CDATA[Building outstanding rich-text experiences with ProseMirror and React]]></title><description><![CDATA[Rich-text editors power knowledge-sharing and content creation on the web.]]></description><link>https://www.brunoscheufler.com/p/2024-02-11-building-outstanding-rich-text-experiences-with-prosemirror-and-react</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-02-11-building-outstanding-rich-text-experiences-with-prosemirror-and-react</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 11 Feb 2024 18:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Rich-text editors power knowledge-sharing and content creation on the web. Back in the day, WYSIWYG editors for building blogs and submitting comments on websites were all the rage. Nowadays, products like Notion lead the way in simplifying the editing experience while enabling powerful use cases. If the editor works, we rarely think about what&#8217;s happening under the hood.</p><p>As it turns out, building rich-text editors is surprisingly complex. From modeling the underlying data structures used for storing formatted content to creating an intuitive UX that supports a natural editing experience, there&#8217;s a ton of heavy lifting that makes a well-designed editor feel simple.</p><p>Previously, implementations used various workarounds and hacks, but with the introduction of the <a href="https://developer.mozilla.org/en-US/docs/Web/HTML/Global_attributes/contenteditable">contenteditable</a> attribute, browsers enabled a new class of RichText editors using the DOM to render and manipulate content.</p><p><a href="https://prosemirror.net/">ProseMirror</a> is a low-level set of building blocks to construct RichText editor experiences. Think of it as a thin abstraction layer on top of contenteditable HTML, taking care of schema enforcement, state management, and content rendering.</p><p>ProseMirror powers popular libraries like <a href="https://tiptap.dev/">TipTap</a> (which is used by other libraries like <a href="https://novel.sh/">Novel</a> and <a href="https://www.blocknotejs.org/">BlockNote</a>). I believe a big part of its popularity comes from being intentionally unopinionated to create a solid foundation for customizable editing experiences.</p><p>In most cases, I recommend using a wrapper like TipTap simply because it&#8217;s easier to work with, but sometimes, you really need all the power. So let&#8217;s dive into the basic concepts behind ProseMirror.</p><h2>Useful resources</h2><p>Before we get started, I&#8217;ve collected some helpful resources that make working with ProseMirror infinitely easier.</p><p>First, the <a href="https://chromewebstore.google.com/detail/prosemirror-developer-too/gkgbmhfgcpfnogoeclbaiencdjkefonj">ProseMirror Developer Tools Chrome extension</a> allows you to view the internal state of any ProseMirror instance detected on a webpage, including the document content (in a tree-like view, more on that later) with its history, plugins, and schema.</p><p>When you want to understand specific parts of ProseMirror, the official <a href="https://prosemirror.net/docs/guide/">guide</a> is an amazing starting point. The <a href="https://prosemirror.net/examples/">examples</a> section also shows what you can build with ProseMirror, including Markdown support, collaborative editing, and text linting.</p><h2>The ProseMirror Document Model</h2><p>ProseMirror is organized around <a href="https://prosemirror.net/docs/guide/#doc">documents</a> to represent content. A document <a href="https://prosemirror.net/docs/ref/#model.Fragment">contains</a> block <a href="https://prosemirror.net/docs/ref/#model.Node">nodes</a>. A block node may be a paragraph, image, list, or other element. Instead of nesting content in a tree structure, however, ProseMirror attempts to lift up inline text content as far as possible, leading to a very shallow structure.</p><p>Imagine we&#8217;re writing the following text:</p><pre><code>&lt;p&gt;Hello &lt;b&gt;World&lt;/b&gt;&lt;/p&gt;</code></pre><p>You may expect the document to look like the following:</p><pre><code>- node: paragraph
  - text: Hello
  - node: bold
    - text: World</code></pre><p>This is not what ProseMirror does, however.</p><pre><code>- node: paragraph
&#9;- text: Hello
&#9;- text: World (bold)</code></pre><p><em><a href="https://prosemirror.net/docs/ref/#model.Mark">Marks</a></em> like bold, emphasis, inline code, links, and more are not nested but attached to the text.</p><p>Flattening the data structure allows ProseMirror to use positioning within a paragraph for more natural text selection and editing operations without tree manipulation. Furthermore, each document has only one valid representation:</p><blockquote><p>Adjacent text nodes with the same set of marks are always combined together, and empty text nodes are not allowed. The order in which marks appear is specified by the <a href="https://prosemirror.net/docs/guide/#schema">schema</a>, too.</p></blockquote><p>If you&#8217;re venturing into the true depths of ProseMirror, nodes have different property combinations:</p><blockquote><p>A typical <code>"paragraph"</code> node will be a textblock, whereas a blockquote might be a block element whose content consists of other blocks. Text, hard breaks, and inline images are inline leaf nodes, and a horizontal rule node would be an example of a block leaf node.</p></blockquote><h2>Defining the Document Schema</h2><p>Now that we&#8217;ve learned how documents are constructed, let&#8217;s continue with schemas. If you&#8217;re building a product like Notion, you might want to allow your documents or pages to include specific blocks like paragraphs, images, code, and more.</p><p>ProseMirror determines whether a document is valid by checking against a predefined <a href="https://prosemirror.net/docs/guide/#schema">schema</a>. This is defined similarly to parser grammar, as you get to define which nodes exist in the system and what they may contain:</p><pre><code>const trivialSchema = new Schema({
  nodes: {
    doc: {content: "paragraph+"},
    paragraph: {content: "text*"},
    text: {inline: true},
    /* ... and so on */
  }
})</code></pre><p>In the example snippet above, a document may include one or more paragraphs, while a paragraph may include zero or more inline text nodes. Schemas can represent more complex data models like <a href="https://github.com/ProseMirror/prosemirror-markdown/blob/master/src/schema.ts">Markdown</a>. If you want to add custom node types like interactive blocks, you need to register the nodes in your schema first.</p><p>In addition to handling validation, the schema also defines how ProseMirror renders your document <a href="https://prosemirror.net/docs/guide/#schema.serialization_and_parsing">in the DOM</a> and how DOM content is parsed back into a node (relevant for enabling copy-and-paste support).</p><h2>Creating an initial state</h2><p>Whenever you create an editor instance, ProseMirror has to keep track of the current content, selection, plugins, and other data, all of which is wrapped up in the <a href="https://prosemirror.net/docs/guide/#state">editor state</a>. Creating the initial state for ProseMirror to use is as simple as</p><pre><code>import {schema} from "prosemirror-schema-basic"
import {EditorState} from "prosemirror-state"

let state = EditorState.create({schema})</code></pre><p>If you&#8217;re editing an existing document, you can pass the previous document right into the <code>create</code> call.</p><h2>Rendering the editor</h2><p>You could use ProseMirror completely headless and create your own view layer on top, but luckily, you don&#8217;t have to. Create an instance of <a href="https://prosemirror.net/docs/ref/#view.EditorView">EditorView</a> by passing the DOM element to render the editor in as well as the initial state and ProseMirror will render the current document.</p><pre><code>let view = new EditorView(document.body, {state})</code></pre><p>This works great when you&#8217;re not running in some other framework like React. Let&#8217;s explore how we can wire up ProseMirror to work together with React.</p><h2>Using React</h2><p>When you&#8217;re building a React application, you&#8217;ll likely want to use React state management primitives like hooks. You might pass the current document content into the editor and expect updates to be passed back up. The NY Times has published a ProseMirror <a href="https://github.com/nytimes/react-prosemirror">wrapper library</a> for React that bridges the two worlds nicely.</p><p>To render the ProseMirror editor within React, simply render the <code>&lt;ProseMirror/&gt;</code> component:</p><pre><code>import { ProseMirror } from "@nytimes/react-prosemirror";

const initialState = ...

export function ProseMirrorEditor() {
  const [mount, setMount] = useState&lt;HTMLElement | null&gt;(null);

  return (
    &lt;ProseMirror mount={mount} defaultState={initialState}&gt;
      &lt;div ref={setMount} /&gt;
    &lt;/ProseMirror&gt;
  );
}</code></pre><p>If you want to lift up the state and manage it yourself, that works too:</p><pre><code>export function ProseMirrorEditor() {
  const [mount, setMount] = useState&lt;HTMLElement | null&gt;(null);
  const [state, setState] = useState(EditorState.create({ schema }));

  return (
    &lt;ProseMirror
      mount={mount}
      state={state}
      dispatchTransaction={(tr) =&gt; {
        setState((s) =&gt; s.apply(tr));
      }}
    &gt;
      &lt;div ref={setMount} /&gt;
    &lt;/ProseMirror&gt;
  );
}</code></pre><p>Internally, the <code>&lt;ProseMirror/&gt;</code> component will create, mount, and manage a ProseMirror EditorView instance through <a href="https://github.com/nytimes/react-prosemirror?tab=readme-ov-file#useeditorview"><code>useEditorView</code></a>.</p><h2>Bonus: Node Views</h2><p>I hear you: All this is great, but can we customize nodes like paragraphs using React? As always, there is more than one solution. You could change the styles using CSS for the entire editor. But if you need to use anything from the React component tree (e.g. user preferences), you&#8217;ll need something different. Unfortunately, the editor view breaks out of the React paradigm: Everything up until the ProseMirror component lives within React, but the EditorView directly manipulates the DOM.</p><p>Fortunately, <a href="https://prosemirror.net/docs/guide/#view.node_views">node views</a> allow you to take back control of the rendering process for a node. And with <a href="https://github.com/nytimes/react-prosemirror?tab=readme-ov-file#usenodeviews">react-prosemirror</a>, node views render React components using <a href="https://react.dev/reference/react-dom/createPortal">Portals</a>.</p><p>Let&#8217;s say we want to apply TailwindCSS classes to every paragraph node. For this, we need to create a node view</p><pre><code>import {
  useNodeViews,
  useEditorEventCallback,
  NodeViewComponentProps,
  react,
} from "@nytimes/react-prosemirror";
import { EditorState } from "prosemirror-state";
import { schema } from "prosemirror-schema-basic";

// Paragraph is more or less a normal React component, taking and rendering
// its children. The actual children will be constructed by ProseMirror and
// passed in here. Take a look at the NodeViewComponentProps type to
// see what other props will be passed to NodeView components.
function Paragraph({ children }: NodeViewComponentProps) {
  const onClick = useEditorEventCallback((view) =&gt; view.dispatch(whatever));
  return &lt;p onClick={onClick}&gt;{children}&lt;/p&gt;;
}

// Make sure that your ReactNodeViews are defined outside of
// your component, or are properly memoized. ProseMirror will
// teardown and rebuild all NodeViews if the nodeView prop is
// updated, leading to unbounded recursion if this object doesn't
// have a stable reference.
const reactNodeViews = {
  paragraph: () =&gt; ({
    component: Paragraph,
    // We render the Paragraph component itself into a div element
    dom: document.createElement("div"),
    // We render the paragraph node's ProseMirror contents into
    // a span, which will be passed as children to the Paragraph
    // component.
    contentDOM: document.createElement("span"),
  }),
};</code></pre><p>This isn&#8217;t the classic node view that ProseMirror exposes, instead, we&#8217;re rendering a React component. Next, we need to register our custom node views in the editor</p><pre><code>const state = EditorState.create({
  schema,
  // You must add the react plugin if you use
  // the useNodeViews or useNodePos hook.
  plugins: [react()],
});

function ProseMirrorEditor() {
  const { nodeViews, renderNodeViews } = useNodeViews(reactNodeViews);
  const [mount, setMount] = useState&lt;HTMLElement | null&gt;(null);

  return (
    &lt;ProseMirror mount={mount} nodeViews={nodeViews} defaultState={state}&gt;
      &lt;div ref={setMount} /&gt;
      {renderNodeViews()}
    &lt;/ProseMirror&gt;
  );
}</code></pre><p>This way, we can use ProseMirror for a natural rich-text editing experience <em>and</em> add our custom components back into the document. It&#8217;s the best of both worlds!</p><div><hr></div><p>ProseMirror is incredibly powerful and can be used for building best-in-class editors on the web. At <a href="https://codetrail.io/">CodeTrail</a>, we&#8217;re building our core documentation experience using ProseMirror, and I&#8217;ll dive deeper into some practical applications in upcoming posts!</p>]]></content:encoded></item><item><title><![CDATA[Reducing Python Memory Usage by 97%]]></title><description><![CDATA[I&#8217;ve recently had the honor to debug a Python script that was consuming more memory than it takes to run a full-sized LLM locally.]]></description><link>https://www.brunoscheufler.com/p/2024-02-04-reducing-python-memory-usage</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-02-04-reducing-python-memory-usage</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 04 Feb 2024 18:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve recently had the honor to debug a Python script that was consuming more memory than it takes to run a full-sized LLM locally. The purpose of the script wasn&#8217;t anything out of the ordinary, it was reading a CSV file and performing some analytics on it. The CSV contained 10m lines of request logs, including a duration in seconds.</p><p>With some debugging, I quickly tracked down the issue to the loading and preprocessing step. Before running any data analysis, the duration should be converted to milliseconds, as that&#8217;s easier to work with and plot later on.</p><pre><code>machine,duration
1,0.045
2,0.035
3,0.358</code></pre><p>As you can see below, the original author used the built-in CSV reader to stream the data into Python. As a hint, the first mistake I made was not looking into how the data is parsed by Python, because the code does look pretty straightforward.</p><pre><code>from csv import reader

rows = []
with open('./test.csv', 'r') as csv_file:
    csv_reader = reader(csv_file)
    next(csv_reader)  # Skip the header row
    for row in csv_reader:
        rows.append(row)

# convert duration to seconds
rows = [(row[0], row[1] * 1000) for row in rows]</code></pre><p>Running the conversion step blew up the memory usage from around 5 GB to over 130 GB. Unfortunately, I had to terminate the script early because my Mac was swapping extensively and everything started to break down past 100 GB.</p><p>Obviously, something was wrong. I reduced the data set to 10 rows to debug every step without waiting forever. Inspecting the layout of the rows quickly yielded an interesting piece of information: Instead of floating-point data types, the duration (and all other rows for that matter) were read as strings. So rather than efficiently storing the values as 32-bit or 64-bit floats, Python&#8217;s CSV reader didn&#8217;t know how to parse the duration without any additional details and opted for strings which offer more possibilities but are the worst possible option for numerical values.</p><p>Multiplying a string with a number in Python does not (in classic JavaScript fashion) convert the value to a number and run with it like below. See for yourself below:</p><pre><code>$ python3
&gt;&gt;&gt; x = "0.045"
&gt;&gt;&gt; x * 10
'0.0450.0450.0450.0450.0450.0450.0450.0450.0450.045'
&gt;&gt;&gt; x = 0.045
&gt;&gt;&gt; x * 10
0.44999999999999996
&gt;&gt;&gt;

$ node
Welcome to Node.js v20.9.0.
&gt; "0.5" * 10
5</code></pre><p>Ah well, this is wrong on so many levels. Let&#8217;s start and fix it.</p><p>If we wanted to keep this script working as is and don&#8217;t introduce any additional libraries, we should convert the data types as early as possible to prevent running into these issues.</p><pre><code>rows = []
with open('./test.csv', 'r') as csv_file:
    csv_reader = reader(csv_file)
    next(csv_reader)  # Skip the header row
    for row in csv_reader:
        row[1] = float(row[1]) # convert strings to floats
        rows.append(row)</code></pre><p>Adding one line to the input reader will convert the duration to proper floats. This simple change brought memory usage down back to around 5 GB and allowed the script to run to completion.</p><p>Alternatively, we could have used a library like pandas to create a data frame from the CSV immediately:</p><pre><code>import pandas as pd
df = pd.read_csv('./test.csv')
df['duration'] = df['duration'] * 1000</code></pre><p>Not only is this much simpler, it&#8217;s way harder to screw up.</p><div><hr></div><p>You might think &#8220;This will never happen to me&#8221; and &#8220;Oh boy, those were some seriously stupid mistakes&#8221;, but I wonder how many codebases end up in a situation like this. The root cause was manifold:</p><p>First, a lack of context blocked me from spotting the mistakes: I usually write JavaScript so the Python way of converting between data types wasn&#8217;t on my mind. I also used Pandas and other libraries in the past and expected the built-in CSV handling to parse input values similarly.</p><p>Second, the lack of a strict type system and validation layer made it harder to spot unexpected differences in data types. It might be helpful to require users to define their column types ahead of time and raise an exception if unexpected input values are supplied.</p><p>Finally, unit tests on smaller but representative data sets would have surfaced the invalid implementation at the time of writing the code.</p>]]></content:encoded></item><item><title><![CDATA[Scaling Generative AI Adoption with Software Engineering Principles]]></title><description><![CDATA[Now that everyone and their dog have tried out Generative AI, teams around the globe are wondering how they can incorporate LLMs into their products for an upcoming fundraising effort or, hopefully, to create a great experience for their customers and solve some real problems.]]></description><link>https://www.brunoscheufler.com/p/2024-01-28-scaling-generative-ai-adoption-with-software-engineering-principles</link><guid isPermaLink="false">https://www.brunoscheufler.com/p/2024-01-28-scaling-generative-ai-adoption-with-software-engineering-principles</guid><dc:creator><![CDATA[Bruno Scheufler]]></dc:creator><pubDate>Sun, 28 Jan 2024 18:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!t_Dh!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a4b6eb-7eb2-4120-a949-7e83d0f6060c_2330x2330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Now that everyone and their dog have tried out Generative AI, teams around the globe are wondering how they can incorporate LLMs into their products for an upcoming fundraising effort or, hopefully, to create a great experience for their customers and solve some real problems. In recent months, <a href="https://www.intel.com/content/www/us/en/newsroom/news/10-per-cent-orgs-launched-genai-solutions-2023.html#gs.47yjvd">countless</a> <a href="https://www.oreilly.com/radar/generative-ai-in-the-enterprise/">reports</a> <a href="https://info.kpmg.us/news-perspectives/technology-innovation/kpmg-generative-ai-2023.html">detailing</a> <a href="https://ai-infrastructure.org/enterprise-generative-ai-adoption-report-aug-2023/">Generative AI</a> <a href="https://mlops.community/surveys/llm/">adoption</a> have presented larger concerns about getting this new technology rolled out in production.</p><p>As a software engineer, I&#8217;ve been following the formation of a vibrant ecosystem of projects around LLMs. As a co-founder of <a href="https://gradientsandgrit.com/">Gradients &amp; Grit</a>, I&#8217;ve explored production-ready ways of integrating LLMs into existing products.</p><p>A lot of progress has been made in the open-source model landscape. Yet, before building system-critical functionality on top of LLMs, teams need to address some underlying issues. In the following, I&#8217;ll explore the core areas holding back LLM adoption in production environments.</p><h2>Interpretability, Reliability, Consistency</h2><p>In the past, it was easy to design systems and trace outputs back to inputs. You could evaluate systems with requirements and compare different implementations. With machine learning and especially deep learning, the connection between inputs and outputs has gotten fuzzy by design. Distributed systems sit at every level, adding randomness and non-determinism to the mix.</p><p>With LLMs, differences in architecture, training data, inference, and parameters supplied in each request can lead to wildly different responses, without accounting for hallucinations. When you&#8217;re consuming a model through an API, there&#8217;s no guarantee that you receive a matching response for the same request twice in a row.</p><p>This creates complexity throughout the entire software engineering lifecycle, from development to production. For starters, it&#8217;s hard to nail the right prompt (one metric could be the best performance for the lowest token usage) for a use case. This prompt isn&#8217;t portable as you can&#8217;t switch the underlying model and expect similar results. You can&#8217;t even expect the same model to yield the same results in the future. With open-source models, it should be easier to pin your system to a specific model version, which is great for consistency. Similar to regular dependencies, you&#8217;ll have to update or at least re-visit prompts once you update the model.</p><p>If you&#8217;re self-hosting the model, you&#8217;ll need to worry about serving it in production, and handling inference on a cluster of machines. This is distributed computing with near real-time requirements, which has been a challenge for every other use case in the past. If you&#8217;re opting for a managed solution, you can offload some complexity at the cost of giving up control and your data. Later on, we&#8217;ll see that the latter is a relevant blocker for enterprise customers and government institutions.</p><p>Of course, in traditional software engineering, you had to spend time to get to a good solution while balancing trade-offs, too. Yet, the degree of fuzziness or lack of guarantees with LLMs and machine learning in general has to be accounted for. You need to be aware of the system&#8217;s biases leading to unexpected consequences for a wrong prediction. This is why most products have incorporated AI in nice-to-have features with human supervision. Running LLMs in the background without any safeguards may be an expensive mistake.</p><p>This is not to say that traditional software isn&#8217;t prone to bugs that can <a href="https://www.bbc.com/news/business-56718036">destroy livelihoods</a>, it&#8217;s just harder to test notoriously unpredictable systems.</p><p>Another interesting facet of working with LLMs is that you&#8217;re switching between strict and fuzzy logic when crossing boundaries between code and prompts. LLMs accept a modicum of inconsistency when it comes to formatting prompts, so chaining prompts works better than expected. Factual inaccuracy in one response might lead to a game of telephone when passed along a chain of LLMs, though.</p><p>Unfortunately, it&#8217;s hard to evaluate correctness once you serialize the result back into code. Forcing models to output a well-defined format like JSON has shown to help, and support for steerability is getting better with newer models.</p><h2>Cost-Effectiveness</h2><p>Let&#8217;s assume we&#8217;ve found a suitable use case for Generative AI and created a prototype for it. How expensive will it be to run this in production? Since OpenAI is subsidized by Microsoft, they can afford to operate at a loss throughout the near future, but unless you&#8217;ve raised a comfortable round recently, you might need to run a tight ship.</p><p>Running your own hardware is expensive. Hosted inference solutions scale with demand, which is great. Tokens can be cheap or expensive, based on the model&#8217;s capabilities. Some tasks need complex reasoning, while others can be done with less. Some prompts need a lot of details, some don&#8217;t. Figuring out the right balance requires a certain playfulness early on in the development process. Balancing different model characteristics can determine how feasible the business plan is down the road.</p><p>It&#8217;s important to understand the costs you will incur at different levels of scale. If it&#8217;s impossible to achieve positive unit economics, you&#8217;ll have a hard time in the best possible scenario. The bigger the role AI plays in delivering your product, the more you need to match your pricing. Traditional compute resources have become commoditized and allowed operating on incredibly high margins. Quantization and other efficiency improvements for models will eventually drive resource prices down.</p><h2>Operational Knowledge</h2><p>All AI-enabled products I&#8217;ve recently worked on sounded simple in principle and required a surprising amount of hardening and preparing for production readiness before launching. Knowing the necessary measures to run and scale a system is as important as designing the product itself.</p><p>It&#8217;s still early enough for operational knowledge to be built and iterated on for the first couple of times. There are few absolute truths and most teams are learning on a daily basis. Some voices criticized that prompt engineering is often lacking the engineering part.</p><p>As with any new paradigm, teams will slowly try out Generative AI and iterate until they&#8217;re confident about delivering a high-quality experience to their customers. It may make sense to build internal tools first and start in the background, then expand to adding AI features to the product. Then, some companies are working at the cutting edge and building their entire business on top of Generative AI.</p><p>I&#8217;m convinced that the community will share their knowledge and work together so that every team can eventually benefit from the progress made in the biggest companies.</p><h2>Tooling</h2><p>While it&#8217;s easy to put together a demo application consuming an API, building production-ready AI-enabled products requires many building blocks. For the same reasons that operational knowledge is still being created, solid tools for building LLM applications are yet to be created.</p><p>From tailoring prompts to analyzing model usage, tracking costs, monitoring errors, and evaluating consistency, the entire LLM lifecycle needs to be covered by the LLMOps movement. It&#8217;s already really easy to consume open-source models through APIs and hosted inference endpoints. Frameworks like Haystack and langchain make it easier to bring your own data, and I imagine more will follow in the months to come.</p><h2>Enterprise: Safety and Data Governance</h2><p>Depending on your customers, offering a product isn&#8217;t as straightforward as plugging into the OpenAI APIs. When you need to guarantee that customer data never leaves the jurisdiction, you&#8217;ll have to resort to regional service providers or resort to hosting models yourself. The latter is quite expensive, but if you&#8217;re aiming for government contracts, trust is an important currency.</p><p>The easier it becomes to run models with a smaller resource footprint, the easier it gets to run inference at the edge where it&#8217;s needed. Moving all components closer together also decreases latency, so customers will enjoy a snappier experience.</p><p>Most current service providers are US-focused and don&#8217;t sufficiently cover other geographic areas and jurisdictions. There&#8217;s a reason why Germany is seeing a ton of startups offering the same services available abroad for a GDPR-conscious enterprise sector. This may not be the best moat, but it&#8217;s certainly a market waiting to be served.</p><div><hr></div><p>I&#8217;m hardly the <a href="https://huyenchip.com/2023/04/11/llm-engineering.html">first</a> <a href="https://mitchellh.com/writing/prompt-engineering-vs-blind-prompting">person</a> who&#8217;s been thinking about running Generative AI in production. There&#8217;s a lot of hype and people expect LLMs to play a huge role in the years to come. To make this happen, we need to ensure that systems are reliable, consistent, cost-effective, and safe. Coming up with features to solve problems is one part of the solution, but proper execution requires solid software engineering. LLMs aren&#8217;t just used by researchers in labs anymore and the ecosystem has to adapt to this new reality.</p>]]></content:encoded></item></channel></rss>