• 42 Posts
  • 375 Comments
Joined 3 years ago
cake
Cake day: June 11th, 2023

help-circle
  • So my main point is to stop asking whether and how code review in its current form can be saved, but to have an open conversation about what we are trying to achieve here and the costs and benefits of alternative models.

    What would these other models be? The article assumes LLM-generated code leads to the loss of current review processes, which they call ‘modern reviews’.

    I feel like my team does and will continue to do reviews for the same reasons and with the same gains as before. LLM may generate code, but I expect my colleagues to take ownership of whatever they produce, to understand it, take responsibility, and describe it. (Which is an issue/gap for some, but that was an issue even before generating code.)





  • 00:00 Intro
    00:48 Why write a compiler in JavaScript
    07:29 Why rewrite the compiler in Go
    14:49 LLMs for large migrations
    20:12 Why Javascript is so popular
    26:32 Why ever use Javascript on the backend
    32:59 What it takes to build a programming language
    37:06 Will there be fewer languages in 10 years
    42:57 Hands on engineering vs delegation
    49:14 Why fast tooling matters more now
    51:16 AI software engineering predictions
    58:52 The most technically challenging work
    01:02:04 Top book recommendation
    01:03:50 Advice for his younger self
    01:05:00 Outro

    Just summarizing the title teasers:

    10x faster by switching from JavaScript to Golang. Rust was not feasible because it does not have a GC/does not support circular dependencies, and they have many of those, and did not want a rewrite but a minimal effort switch.

    AI may produce a lot more code, but what code does “90% of code” refer to. AI is not capable of novel or complex stuff, you still need humans in the loop. Juniors remain significant to train them into seniors.






  • Simple agents won’t understand git. They can make web requests. If they want to check for reference code/source code, I assume they would request the rendered html.

    I assume the frequency comes from many people asking various things, and the agents in the background pulling this data.

    I’m not sure whether the scale-to-load ratio is plausible, because I lack the numbers, but it doesn’t seem implausible to me that various agents for various prompts and tasks from various providers for many people repeatedly make these requests to a degree that significantly impacts the hoster.




  • if I was going to expand my horizons

    Python is well liked and beloved, but it never captivated me or felt intuitive.

    Writing magic init logic feels like a hacky workaround, like putting frameworks and concepts around what is, at its core, a scripting language. If you’re committed to it, maybe you’ll get past that barrier and find it’s good, I don’t know, I never got there. From outside, and with experience in other languages/ecosystems, it just never felt correct to me.

    Have you considered alternatives? For example, C#/dotnet?

    Just my experience/opinion. As I said, it seems well beliked by others. I can’t speak for their views or experiences.



  • TL;DR: we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.

    I’ve read similar and even worse before. Probably from Codeberg, or maybe a smaller host of Forgejo or GitLab.

    The blame html pages were repeatedly being requested. Which is even worse and even less plausibly useful than a commit html page.


  • They’re not scraping to cache or store, they’re operating as an agent - scraping or single user requests.

    Which is obviously bad and damaging, especially on their scale and on repeatedly fetched websites that they could be caching.

    Google indexed the entire web. It’s baffling that such indexing is not the norm on these huge providers.

    Just my interpretation anyway.







  • Very interesting. It certainly shows the potential a replacement of the very old ELF could offer. It even looks very plausibly doable. Now it just needs to get all distros on board. :P

    I’m not sure I understood the jump from single file SELF to userland-wide.

    Can the SQL be precompiled into binary for faster lookups? I imagine there’s not much variance in the standard queries. That could offer some speedup, skipping the query analyzer/tokenizer/interpreter.

    I’m not sure I followed how the LD_PRELOAD equivalent works now.

    What are disadvantages of using SELF? (Assuming a full OS switch to SELF. Mixed is always additional burden.)