commit 340a6627d52a77923c71aacc92db0d378dd18dd9
parent 15e2e748d44d86dddfdf468710a98eb665dfc8aa
Author: Andrew Laack <andrew@laack.co>
Date: Fri, 18 Sep 2026 00:49:01 -0500
Added assets dir for benchmarking information
Diffstat:
11 files changed, 975 insertions(+), 0 deletions(-)
diff --git a/assets/abg/benchmarking/bench.sh b/assets/abg/benchmarking/bench.sh
@@ -0,0 +1,5 @@
+while [ 1 ]; do
+ /usr/bin/time -o out -f '%S,%U,%e' ./abg.out -s 0 --edges 100000 --vertices 10000 >/dev/null 2>&1
+ cat out | tee -a benchmarking/render_smart_push_v_10000_e_100000_s_0/out.csv
+ sleep 1
+done
diff --git a/assets/abg/benchmarking/micro-sampling/rlib-no-render.txt b/assets/abg/benchmarking/micro-sampling/rlib-no-render.txt
@@ -0,0 +1,47 @@
+render time us: 8310
+prim step us: 6
+render time us: 179
+prim step us: 27
+render time us: 315
+prim step us: 7
+render time us: 198
+prim step us: 8
+render time us: 210
+prim step us: 7
+render time us: 487
+prim step us: 7
+render time us: 243
+prim step us: 6
+render time us: 201
+prim step us: 8
+render time us: 261
+prim step us: 21
+render time us: 200
+prim step us: 8
+render time us: 549
+prim step us: 6
+render time us: 202
+prim step us: 11
+render time us: 248
+prim step us: 7
+render time us: 177
+prim step us: 7
+render time us: 240
+prim step us: 6
+render time us: 833
+prim step us: 9
+render time us: 280
+prim step us: 36
+render time us: 316
+prim step us: 7
+render time us: 299
+prim step us: 8
+render time us: 245
+prim step us: 10
+render time us: 1085
+prim step us: 10
+render time us: 747
+prim step us: 8
+render time us: 235
+prim step us: 11
+render time us: 373
diff --git a/assets/abg/benchmarking/micro-sampling/rlib-no-vertex.txt b/assets/abg/benchmarking/micro-sampling/rlib-no-vertex.txt
@@ -0,0 +1,35 @@
+render time us: 10723
+prim step us: 11
+render time us: 1657
+prim step us: 9
+render time us: 1276
+prim step us: 7
+render time us: 1031
+prim step us: 22
+render time us: 900
+prim step us: 7
+render time us: 992
+prim step us: 7
+render time us: 1502
+prim step us: 10
+render time us: 1056
+prim step us: 7
+render time us: 1037
+prim step us: 6
+render time us: 1636
+prim step us: 8
+render time us: 977
+prim step us: 6
+render time us: 1603
+prim step us: 7
+render time us: 1084
+prim step us: 10
+render time us: 1001
+prim step us: 6
+render time us: 1763
+prim step us: 8
+render time us: 1119
+prim step us: 6
+render time us: 960
+prim step us: 10
+render time us: 1089
diff --git a/assets/abg/benchmarking/micro-sampling/rlib-optimized-edge.txt b/assets/abg/benchmarking/micro-sampling/rlib-optimized-edge.txt
@@ -0,0 +1,45 @@
+render time us: 12067
+prim step us: 9
+render time us: 2389
+prim step us: 11
+render time us: 3500
+prim step us: 6
+render time us: 2212
+prim step us: 12
+render time us: 2397
+prim step us: 9
+render time us: 3252
+prim step us: 17
+render time us: 2355
+prim step us: 8
+render time us: 3597
+prim step us: 8
+render time us: 3294
+prim step us: 6
+render time us: 910
+prim step us: 9
+render time us: 2810
+prim step us: 10
+render time us: 2219
+prim step us: 10
+render time us: 2504
+prim step us: 10
+render time us: 3424
+prim step us: 7
+render time us: 2268
+prim step us: 7
+render time us: 3821
+prim step us: 10
+render time us: 3005
+prim step us: 11
+render time us: 2118
+prim step us: 10
+render time us: 2417
+prim step us: 7
+render time us: 2314
+prim step us: 8
+render time us: 2934
+prim step us: 9
+render time us: 3074
+prim step us: 9
+render time us: 2458
diff --git a/assets/abg/benchmarking/micro-sampling/rlib-prim.txt b/assets/abg/benchmarking/micro-sampling/rlib-prim.txt
@@ -0,0 +1,59 @@
+render time us: 6640
+prim step us: 7
+render time us: 1946
+prim step us: 7
+render time us: 2985
+prim step us: 11
+render time us: 2229
+prim step us: 6
+render time us: 1882
+prim step us: 10
+render time us: 1970
+prim step us: 7
+render time us: 2082
+prim step us: 8
+render time us: 3345
+prim step us: 10
+render time us: 1873
+prim step us: 7
+render time us: 1868
+prim step us: 5
+render time us: 794
+prim step us: 8
+render time us: 2027
+prim step us: 10
+render time us: 2774
+prim step us: 9
+render time us: 1970
+prim step us: 7
+render time us: 2015
+prim step us: 12
+render time us: 1879
+prim step us: 7
+render time us: 1917
+prim step us: 9
+render time us: 2304
+prim step us: 16
+render time us: 1916
+prim step us: 10
+render time us: 1845
+prim step us: 10
+render time us: 3320
+prim step us: 11
+render time us: 1958
+prim step us: 12
+render time us: 3531
+prim step us: 15
+render time us: 2111
+prim step us: 12
+render time us: 1847
+prim step us: 10
+render time us: 1982
+prim step us: 11
+render time us: 1945
+prim step us: 11
+render time us: 2671
+prim step us: 10
+render time us: 2017
+prim step us: 14
+render time us: 1936
diff --git a/posts/wip/default-cracked.md b/posts/wip/default-cracked.md
@@ -0,0 +1,3 @@
+# Default: Cracked
+
+Language models aren't default cracked. Be default cracked; stand in their face and spit in it.
diff --git a/posts/wip/do-tools-actually-matter.md b/posts/wip/do-tools-actually-matter.md
@@ -0,0 +1,7 @@
+# Do Tools Actually Matter
+
+I think about fuzzing / pbt and cool outcomes. that's great. Is this just on the edge? Are the kinds of ppl who use these things the kinds of people who would find these issues even if they had shit tools? I was trying to find a snapshot based testing tool in C++ to validate serialized data models and realized I don't really give a fuck and can just use asserts, one write string to file function, one read file to string.
+
+## Does it matter
+
+I think vet kinda matters because there are optimizations we performed to find issues. So in a sense, that's better than just using a LM and saying find shit. In a similar vein, if we compare something using a better complexity heuristic, like LEOPARD-V, it'll be a better (fuzzer) than if it was just using some sort of naive code coverage metric, assuming consistent runtimes and non-exhaustive search.
diff --git a/posts/wip/greyscale-shader.md b/posts/wip/greyscale-shader.md
@@ -0,0 +1 @@
+greyscale shader is cool. made eyes hurt once? will happen again?
diff --git a/posts/wip/to-make-a-cool-background.md b/posts/wip/to-make-a-cool-background.md
@@ -0,0 +1,724 @@
+# To Make a Cool Background
+
+I wrote a visualization program for prim's algorithm that induces a minimum spanning tree (MST) on a procedurally generated graph. How does one go about making this their background and screensaver with X11 / dwm?
+
+## MST
+
+IMPLEMENTATION DETAILS
+
+## Background making
+
+https://github.com/python-xlib/python-xlib
+
+rather shit; also, this isn't faster than using feh from a screenshot of the pygame stuff
+
+> os.environ["SDL_VIDEODRIVER"] = "dummy"
+
+why not just use screenshots in a shared location?
+
+indeed, why not?
+
+## fuck you and your bullshit.
+
+Benchmarking
+
+## c -> c++
+
+realized I needed a heap.
+
+## bmp vs png vs jpg
+
+bmp is crazy fucking fast (benchmark)
+
+(looked in the code for supported file formats, looked in feh -> lib(whatever) for support)
+
+note larger file size.
+
+this is interesting, look further.
+
+### patched upstream xl
+
+### x11
+
+x11 is weird with this because they are kind hacky in that they just put up a window and take away your inputs, forcing them to go to that window.
+
+(should learn more about this)
+
+### Benchmarking
+
+uhhh... is it actually that fast?
+
+wait; did this just optimize all of my code?
+
+Well yes, but actually no.
+
+adding:
+
+std::cout << g.toString() << std::endl;
+
+gives ~ the same perf (minus the amount of time it takes to send so much data to stdout)
+
+(this would require the full traversal because toString is dependent upon prior prim algo steps being performed otherwise the traversal status for edges and vertices would be incorrect).
+
+(fixed a bug in the python code slowing it down)
+
+see no_render_2_fixed_itr.
+
+## C++ benchmarking
+
+330 ms for init stuff w/ graph
+1405 ms for iteration
+
+broken further down:
+
+average iteration time is ~0ms, reaching ~20ms for later stages. Over time the percentage of valid edges declines.
+
+added getUnvisitedEdges function to return only non-traversed edges
+
+Init time: 332
+Loop time: 1344
+Init time: 327
+Loop time: 1350
+Init time: 322
+Loop time: 1374
+
+still has a similar issue as before; degredation in perf towards the end bc things are pushed earlier on and then get invalidated later on.
+
+that's basically all I could squeeze out without changes to the algo to be eager instead of lazy wrt tracking vertices. Basically, what's the shortest distance to a given vertex instead of what's the edge that we have that's shortest.
+
+such a refactor though is likely irrelevant right now due to the added cost of
+
+ope; realized my function for get edges with unvisited vertices above was wrong. I was just not pushing visited edges, but we really care about vertices that haven't been visited so refactored the edge generation and stuff to get a stable ref to v2index and then push only if that hasn't been visited.
+
+Init time: 389
+Loop time: 706
+Init time: 315
+Loop time: 671
+Init time: 337
+Loop time: 692
+
+now we're cooking with gas.
+
+tests now fail. bc snapshots. updated and everything seems right again.
+
+actually the stable things isn't really necessary as this approach that's unstable is functionally the same, but safer in case there are changes to the IR of edges.
+
+Init time: 334
+Loop time: 694
+Init time: 321
+Loop time: 668
+Init time: 323
+Loop time: 701
+
+now we find this:
+
+(see benchmarking/no_render_v_100000_e_1000000_s_0/out1.csv)
+
+same optimization performed for python code so that's then down to 8.6 seconds ~~~ see andrew@deepthought:no_render_3_fixed_itr_fixed_check$
+
+at this point it's worth worrying about the return for the get untraversed edges for vertices not traversed. after that:
+
+Init time: 314
+Loop time: 678
+Init time: 312
+Loop time: 671
+Init time: 311
+Loop time: 680
+
+and with the optimization for edge ordering:
+
+Init time: 311
+Loop time: 659
+Init time: 333
+Loop time: 673
+Init time: 320
+Loop time: 709
+Init time: 311
+Loop time: 674
+
+this is still too fucking slow.
+
+what's making the graph induction so incredibly slow?
+
+just graph creation:
+
+graph create: 317
+Loop time: 691
+graph create: 314
+Loop time: 657
+graph create: 311
+Loop time: 665
+
+reserving instead of pushing:
+
+graph create: 314
+Loop time: 636
+graph create: 324
+Loop time: 688
+graph create: 315
+Loop time: 693
+
+replaced map with vector where we were using indices to index into the map (dumb):
+
+graph create: 174
+Loop time: 654
+graph create: 179
+Loop time: 610
+graph create: 184
+Loop time: 618
+graph create: 185
+Loop time: 622
+
+ok, let's now run the full benchmarks again:
+
+benchmarking/no_render_only_grab_unvisited_edges_vector_instead_of_map_v_100000_e_1000000_s_0/out1.csv
+
+0.04,0.79,0.84
+0.03,0.79,0.84
+0.04,0.77,0.82
+
+this was kind of just stalling though becaus I'm out of my depth as it relates to rendering, the larger perf issue.
+
+back to that.
+
+here are where this stands w/ 1000 v and 10_000 edges
+
+prior was ~34s per iteration (final col)
+
+this is basically the same:
+
+benchmarking/better_render_v_1000_e_10000_s_0/out1.csv
+
+(snippet from above file)
+
+0.26,10.71,34.22
+0.29,10.64,33.24
+0.31,10.62,33.57
+0.28,10.53,33.56
+
+fine. we'll fix the raylib shit. (queue montage)
+
+make debug-build && time ./abg.out -s 0 --edges 2000 --vertices 200
+
+real 0m1.948s
+user 0m0.529s
+sys 0m0.060s
+
+~20 milliseconds per render call w/ make debug-build && time ./abg.out -s 0 --edges 20000 --vertices 2000
+
+Total single iteration cost: 65
+Amount of time calling render: 18
+Total single iteration cost: 65
+Amount of time calling render: 18
+Total single iteration cost: 65
+Amount of time calling render: 18
+Total single iteration cost: 65
+Amount of time calling render: 18
+Total single iteration cost: 65
+Amount of time calling render: 19
+Total single iteration cost: 65
+Amount of time calling render: 18
+Total single iteration cost: 65
+Amount of time calling render: 18
+Total single iteration cost: 65
+
+huh; what's the other iteration cost?
+
+
+ while (!WindowShouldClose() && toVisit.size() != 0) {
+ auto startLoop = std::chrono::steady_clock::now();
+ BeginDrawing();
+ ClearBackground(BLACK);
+ auto startRender = std::chrono::steady_clock::now();
+ g.render();
+ auto endRender = std::chrono::steady_clock::now();
+ auto diff = std::chrono::duration_cast<std::chrono::milliseconds>(endRender - startRender);
+ std::cout << "Amount of time calling render: " << diff.count() << std::endl;
+ EndDrawing();
+
+ // since we wait sleepTime here, the bg render has a render delta of
+ // at minimum sleepTime when switching tags in dwm, this is rather
+ // annoying because the screen doesn't repaint until the sleep time
+ // passes, which results in artifacts on screen.
+
+ // despite this, calling render a lot of times is rather intensive
+ // (at least on my hardware) and so this tradeoff is accepted for
+ // now, unless there's a simple approach that allows for preemption
+
+ //usleep((int)(sleepTime * 1000000));
+
+ oneStepPrim(toVisit, visitedIndices, g);
+ auto endLoop = std::chrono::steady_clock::now();
+ auto loopDiff = std::chrono::duration_cast<std::chrono::milliseconds>(endLoop - startLoop);
+ std::cout << "Total single iteration cost: " << loopDiff.count() << std::endl;
+
+ }
+
+so it's somewhere in that first block; in the render area. the actual computation cost at the end is inconsequential.
+
+begindrawing is ~free.
+
+clear background is ~free.
+
+~46 of 65ms are spent on enddrawing.
+
+so two things:
+
+19ms on render
+46ms on enddrawing
+
+what about when the numbers are much bigger (100_000 edges, 10_000 vertices)?
+
+Render time: 56
+End Drawing time: 237
+Total single iteration cost: 316
+
+the other 23 ms is the prim step + begin drawing / clear background.
+
+can we just not clear the screen? does that help? nope. doesn't do anything.
+
+target 60 fps?
+
+nope.
+
+ahh; the slowness is being realized later on. while render is cool with me double calling on edges; this is felt when enddrawing is called as the gpu does the rendering and stuff at that point.
+
+what happens if I use the better single time render method?
+
+before:
+
+Render time: 62
+End Drawing time: 244
+Total single iteration cost: 306
+Render time: 60
+End Drawing time: 245
+Total single iteration cost: 305
+Render time: 60
+End Drawing time: 245
+Total single iteration cost: 305
+Render time: 55
+End Drawing time: 252
+Total single iteration cost: 307
+Render time: 60
+End Drawing time: 245
+Total single iteration cost: 305
+Render time: 58
+End Drawing time: 247
+Total single iteration cost: 306
+
+after:
+
+Total single iteration cost: 158
+Render time: 46
+End Drawing time: 112
+Total single iteration cost: 159
+Render time: 46
+End Drawing time: 112
+Total single iteration cost: 159
+Render time: 45
+End Drawing time: 113
+Total single iteration cost: 159
+Render time: 46
+End Drawing time: 111
+Total single iteration cost: 158
+Render time: 46
+End Drawing time: 112
+Total single iteration cost: 158
+Render time: 46
+End Drawing time: 113
+
+hell yeah.
+
+why so slow still?
+
+Like this is only ~100_000 and ~10_000 vertices so like why so slow?
+
+benchmarked now:
+
+benchmarking/simpler_render_v_1000_e_10000_s_0/out1.csv
+0.29,7.77,18.96
+0.27,7.80,19.17
+0.29,7.90,19.17
+0.31,7.76,19.05
+0.29,7.72,18.68
+0.30,7.90,18.56
+0.27,7.78,18.80
+0.27,7.77,18.74
+0.29,7.69,18.77
+0.32,7.67,18.85
+
+alright then. unfortunately, we're only calling the draw function once per vertex and edge now. we could change the rendering to only render connections, this would increase speed a lot, but it also makes the animation look less cool so fuck that.
+
+could we draw in the background? like create a new thread that doesn't block us and then wait until that's done once we get back?
+
+well maybe, but the issue is I don't know how well this handles drawing while there's a queue of messages still being sent.
+
+std::thread t(endDrawing);
+t.join();
+
+vs endDrawing()
+
+result in entirely different outcomes.
+
+the second sort of renders some stuff, but gets totally fucked up and then just black screens.
+
+due to an import? no this shit's just fucked.
+
+opengl affinity.
+
+can prebake the graph though and then only render additions.
+
+after:
+
+0.19,1.77,3.86
+0.17,1.76,3.81
+0.19,1.76,3.82
+0.18,1.75,3.85
+0.18,1.71,3.86
+0.19,1.68,3.83
+0.18,1.72,3.83
+
+holy shit. can we make this faster?
+
+of fucking course we can. track if a traversed state has been rendered and skip it if it has, using this to update our incremental texture. we then render that.
+
+benchmarking/cache_first_use_blank_graph_v_1000_e_10000_s_0/out1.csv
+
+0.17,0.65,3.06
+0.18,0.63,3.05
+0.18,0.66,3.05
+0.18,0.64,3.05
+0.18,0.64,3.04
+0.17,0.63,3.02
+0.18,0.62,3.04
+0.19,0.64,3.05
+0.17,0.65,3.04
+0.17,0.66,3.03
+0.18,0.62,3.04
+0.17,0.63,3.03
+
+this is too boring; let's turn this shit up.
+
+how fast are we relative to the python bloatware webshit?
+
+benchmarking/cache_first_use_blank_graph_v_10000_e_100000_s_0/out1.csv
+
+1.40,28.92,35.76
+1.53,29.30,35.95
+
+(notice the edges and vertices)
+
+vs python:
+
+1.33,10730.71,10772.74
+
+that's only >300x slower for the python code so the average python developer would totally be cool with shipping it. Job done, webshit: built.
+
+alright then. enough messing around. We really should have a queue for items that are to be rendered so we don't have to iterate over all vertices and edges, checking if they have the 'rendered' flag set.
+
+after doing this we find:
+
+benchmarking/track_to_render_v_10000_e_100000_s_0/out1.csv
+
+1.50,3.12,29.43
+1.51,3.20,29.67
+1.57,3.08,30.00
+1.57,3.04,30.07
+
+35.855000000000004 / 29.7925 = 1.2034908114458338
+
+okay, so 20 percent faster. still too fucking slow.
+
+i'm also getting sick of this bs chrono / time shit. time to bring in perf.
+
+ 3.20% 395 abg.out libraylib.so.6.0.0 [.] rlVertex3f
+ 2.08% 338 abg.out abg.out [.] Edge::operator>(Edge const&) const
+ 1.55% 245 abg.out abg.out [.] oneStepPrim(std::priority_queue<Edge, std::vector<Edge, st>
+ 1.41% 255 abg.out libraylib.so.6.0.0 [.] rlDrawRenderBatch
+ 1.22% 225 abg.out libc.so.6 [.] 0x0000000000185dde
+ 0.85% 129 abg.out libraylib.so.6.0.0 [.] DrawCircleSector
+ 0.79% 136 abg.out libc.so.6 [.] pthread_mutex_lock
+ 0.74% 133 abg.out libc.so.6 [.] ioctl
+ 0.64% 115 abg.out libgallium-26.2.2-arch1.1.so [.] 0x00000000007f8138
+ 0.62% 113 abg.out libraylib.so.6.0.0 [.] PollInputEvents
+ 0.57% 91 abg.out libc.so.6 [.] cfree
+ 0.56% 95 abg.out libc.so.6 [.] malloc
+ 0.54% 27 abg.out libc.so.6 [.] 0x00000000001866c9
+ 0.53% 99 abg.out libgallium-26.2.2-arch1.1.so [.] 0x000000000180c45c
+
+alright so raylib is spending a decent amount of time on Vertex3f though not a crazy amount. Comparisons between edges are common; not surprising as we are using a heap with lots of elements towards the end, as we've seen. prim's algorithm generally taking some time, makes sense.
+
+render drawing is rather low, but within reason. This is rather disappointing. It's nice when it's like, no, dude, you spent 99.95% of your time in one function.
+
+looking at the call stack:
+
+|--14.80%--EndTextureMode
+|--58.02%--EndDrawing
+--1.62%--Graph::render()
+
+well. huh. end texture mode is where I add things though the function itself is just clearing the queue.
+
+It's a bit surprising enddrawing is more expensive since it just loads in the texture I've baked.
+
+---
+
+I shall call this ~done. final findings with
+
+rendered below is using vertices = 10_000, edges = 100_000
+
+unrendered below is using vertices = 200_000, edges = 2_000_000
+
+- python (with overdraw mostly fixing + texture map)
+ - rendered
+ - 58.758 seconds
+ - unrendered
+ - 20.253333333333334 seconds
+
+- c++ (with overdraw fixing + texture map)
+ - rendered (I don't know if this is right)
+ - track_to_render_v_10000_e_100000_s_0/out1.csv
+ - unrendered
+ - 2.118 seconds
+
+
+benchmarking/final_no_render_v_200000_e_2000000_s_0/
+
+takeaway:
+
+the rendered python implementation w/ 10_000 v and 100_000 edges using perf:
+
+Performance counter stats for 'python3 prim.py':
+
+ 0 context-switches:u # 0.0 cs/sec cs_per_second
+ 0 cpu-migrations:u # 0.0 migrations/sec migrations_per_second
+ 14,233 page-faults:u # 252.0 faults/sec page_faults_per_second
+ 56,487.57 msec task-clock:u # nan CPUs CPUs_utilized
+ 30,763,483 branch-misses:u # 0.7 % branch_miss_rate (88.90%)
+ 4,305,442,939 branches:u # 76.2 M/sec branch_frequency (88.90%)
+ 128,137,043,750 cpu-cycles:u # 2.3 GHz cycles_frequency (88.90%)
+ 18,736,054,409 instructions:u # 0.1 instructions insn_per_cycle (88.85%)
+ TopdownL1 # 0.1 % tma_backend_bound
+ # 98.9 % tma_bad_speculation (88.87%)
+ # 0.5 % tma_frontend_bound (77.81%)
+ # 0.5 % tma_retiring (88.92%)
+
+ 59.936596362 seconds time elapsed
+
+ 52.628026000 seconds user
+ 1.492521000 seconds sys
+
+gpu constrained though.
+
+the rendered c++ implementation w/ 10_000 v and 100_000 edges using perf:
+
+ Performance counter stats for './abg.out -s 0 --vertices 10000 --edges 100000':
+
+ 0 context-switches:u # 0.0 cs/sec cs_per_second
+ 0 cpu-migrations:u # 0.0 migrations/sec migrations_per_second
+ 9,363 page-faults:u # 1819.4 faults/sec page_faults_per_second
+ 5,146.21 msec task-clock:u # nan CPUs CPUs_utilized
+ 28,144,835 branch-misses:u # 5.9 % branch_miss_rate (89.55%)
+ 473,319,266 branches:u # 92.0 M/sec branch_frequency (88.64%)
+ 3,973,233,282 cpu-cycles:u # 0.8 GHz cycles_frequency (88.55%)
+ 2,925,854,560 instructions:u # 0.7 instructions insn_per_cycle (89.12%)
+ TopdownL1 # 2.7 % tma_backend_bound
+ # 93.6 % tma_bad_speculation (88.66%)
+ # 0.9 % tma_frontend_bound (77.75%)
+ # 2.8 % tma_retiring (88.67%)
+
+ 29.577542431 seconds time elapsed
+
+ 2.925597000 seconds user
+ 2.094917000 seconds sys
+
+32x less instructions than the python version. Neither is cpu constrained, but this is a non-trivial amount of overhead from the python version; c++ version also had 7x higher instructions per cycle.
+
+gpu constrained too.
+
+memory overhead?
+
+c++ rendered sits at ~171mb-175mb. it's not great, hasn't been optimized, def room for improvement.
+
+python rendered sits at ~240mb-248mb
+
+the default invocation method for the c++ implementation I use as my lock screen is running at 160mb. This is rather annoying, but also, it's doing a bunch of rendering stuff so IG that's fine, and this does improve the computational cost in terms of cycles so...
+
+huh... after 2 hours it's chilling at 163mb (TODO).
+
+bc it's always holding onto at least one frame which is 5120x1440 pixels which is ~21mb minimum (assuming r g and b bytes per pixel.)
+
+an eager approach for this would be better, but c++ stl doesn't have an indexed priority queue, so I'll just retcon what I have.
+
+basically, I just want to minimize the useless things I push to the queue. One way to do this is to track the minimum weighted edge with an untraversed vertex and then updating this and only pushing edges with it when they are < that weight.
+
+before:
+
+benchmarking/final_no_render_v_200000_e_2000000_s_0/
+
+0.08,2.03,2.12
+0.09,2.02,2.12
+0.09,2.00,2.11
+0.09,2.00,2.11
+0.08,1.99,2.09
+0.07,1.99,2.08
+0.10,2.00,2.11
+0.09,2.00,2.11
+0.09,2.07,2.18
+0.06,2.08,2.15
+
+
+after:
+
+benchmarking/final_min_no_render_v_200000_e_2000000_s_0/out.csv
+
+0.06,0.97,1.04
+0.06,0.93,1.00
+0.07,0.95,1.03
+0.09,0.92,1.02
+0.08,0.94,1.02
+0.08,0.94,1.03
+0.08,0.94,1.03
+0.06,0.95,1.02
+0.08,0.93,1.02
+0.07,0.95,1.02
+0.07,0.95,1.03
+0.07,0.95,1.03
+0.08,0.95,1.04
+0.09,0.93,1.03
+0.07,0.98,1.06
+0.07,0.95,1.04
+0.08,0.94,1.03
+0.07,0.96,1.04
+0.07,0.95,1.04
+0.07,1.00,1.08
+
+this doesn't really impact time wrt rendered because that is basically all spent on rendering not computation.
+
+do I have a memory leak?
+
+
+
+> time valgrind --tool=memcheck abg
+
+==944657== Process terminating with default action of signal 2 (SIGINT)
+==944657== at 0x50D3952: __syscall_cancel_arch (syscall_cancel.S:56)
+==944657== by 0x5117AEC: internal_syscall_cancel (sysdep-cancel.h:53)
+==944657== by 0x5117AEC: clock_nanosleep@@GLIBC_2.17 (clock_nanosleep.c:48)
+==944657== by 0x5123F26: nanosleep (nanosleep.c:25)
+==944657== by 0x51532E9: usleep (usleep.c:31)
+==944657== by 0x40065B8: main (in /usr/local/bin/abg)
+==944657==
+==944657== HEAP SUMMARY:
+==944657== in use at exit: 11,776,171 bytes in 27,368 blocks
+==944657== total heap usage: 86,922 allocs, 59,554 frees, 29,848,973 bytes allocated
+==944657==
+==944657== LEAK SUMMARY:
+==944657== definitely lost: 0 bytes in 0 blocks
+==944657== indirectly lost: 0 bytes in 0 blocks
+==944657== possibly lost: 6,876,540 bytes in 3,237 blocks
+==944657== still reachable: 4,899,631 bytes in 24,131 blocks
+==944657== suppressed: 0 bytes in 0 blocks
+==944657== Rerun with --leak-check=full to see details of leaked memory
+==944657==
+==944657== For lists of detected and suppressed errors, rerun with: -s
+==944657== ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)
+
+
+real 12m18.560s
+user 0m23.848s
+sys 0m0.969s
+
+andrew@deepthought:~$ time valgrind --tool=memcheck abg
+==947393== Memcheck, a memory error detector
+==947393== Copyright (C) 2002-2024, and GNU GPL'd, by Julian Seward et al.
+==947393== Using Valgrind-3.25.1 and LibVEX; rerun with -h for copyright info
+==947393== Command: abg
+==947393==
+^C==947393==
+==947393== Process terminating with default action of signal 2 (SIGINT)
+==947393== at 0x50D3952: __syscall_cancel_arch (syscall_cancel.S:56)
+==947393== by 0x5117AEC: internal_syscall_cancel (sysdep-cancel.h:53)
+==947393== by 0x5117AEC: clock_nanosleep@@GLIBC_2.17 (clock_nanosleep.c:48)
+==947393== by 0x5123F26: nanosleep (nanosleep.c:25)
+==947393== by 0x51532E9: usleep (usleep.c:31)
+==947393== by 0x40065B8: main (in /usr/local/bin/abg)
+==947393==
+==947393== HEAP SUMMARY:
+==947393== in use at exit: 11,656,499 bytes in 27,404 blocks
+==947393== total heap usage: 64,820 allocs, 37,416 frees, 24,650,054 bytes allocated
+==947393==
+==947393== LEAK SUMMARY:
+==947393== definitely lost: 0 bytes in 0 blocks
+==947393== indirectly lost: 0 bytes in 0 blocks
+==947393== possibly lost: 6,744,916 bytes in 3,235 blocks
+==947393== still reachable: 4,911,583 bytes in 24,169 blocks
+==947393== suppressed: 0 bytes in 0 blocks
+==947393== Rerun with --leak-check=full to see details of leaked memory
+==947393==
+==947393== For lists of detected and suppressed errors, rerun with: -s
+==947393== ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)
+
+
+real 0m53.077s
+user 0m20.121s
+sys 0m0.750s
+
+andrew@deepthought:~$ time valgrind --leak-check=full --tool=memcheck abg -s 0
+
+==948338==
+==948338== LEAK SUMMARY:
+==948338== definitely lost: 0 bytes in 0 blocks
+==948338== indirectly lost: 0 bytes in 0 blocks
+==948338== possibly lost: 6,948,380 bytes in 3,271 blocks
+==948338== still reachable: 4,899,399 bytes in 24,179 blocks
+==948338== suppressed: 0 bytes in 0 blocks
+==948338== Reachable blocks (those to which a pointer was found) are not shown.
+==948338== To see them, rerun with: --leak-check=full --show-leak-kinds=all
+==948338==
+==948338== For lists of detected and suppressed errors, rerun with: -s
+==948338== ERROR SUMMARY: 2838 errors from 2838 contexts (suppressed: 0 from 0)
+
+
+real 21m9.471s
+user 19m30.601s
+sys 1m12.207s
+
+seems like this was just me not closing the window. fixed that.
+
+---
+
+profiling:
+
+start: 121mb (time 5:15pm after running for ~a minute)
+check-in: 121mb (time 8:55pm) -- this is just the default `abg` cmd so not a ton of churn
+stopped: 121mb (9:11pm)
+
+running again w/ s=0:
+
+---
+
+(recompile)
+
+start: 123mb (time 9:13pm)
+end: 123mb (time: 11:00pm)
+
+---
+
+seems good; calling it here. further refactors can be done that'd improve perf, mostly be improving cache utilization by removing unnecessary bits. This is what I'd say:
+
+- It's unlikely there's an optimization for my hardware that'd improve performance by 2x
+- It's unlikely there's an optimization for my hardware that'd improve performance by 5x across the workloads I've evaluated. There are ways to make prim's algorithm much faster on much larger graphs, but given what I care about, this is likely inconsequential.
+- These further optimizations are probably not worth it
+ - They could make this faster but two things:
+ 1. They would decrease code readability and extensibility a non-trivial amount
+ - sometimes worth it
+ 2. The improvements wouldn't be meaningful for average usage
+ - most ppl will be draw constrained. I believe I'm close to what is optimal as it relates to using raylib, without diving straight into opengl
+
+
+---
+
+Pygame vs raylib:
+
+- Not a super great comparison
+- SDL (pygame) will basically always be slower than opengl (raylib)
diff --git a/posts/wip/what-is-a-lockscreen.md b/posts/wip/what-is-a-lockscreen.md
@@ -0,0 +1,46 @@
+# The Humble Lock Screen
+
+Lock screens are used to limit access to a system by requiring a user to perform a ritual. These rituals can be arbitrary so long as the individuals capable of the ritual are the set of individuals you're okay with using the system.
+
+As of writing this, I use [xl](https://github.com/dannyfiresnake/xl) as my lock screen. I started using this because I wanted my awesome background to display when my screen was locked, and after learning how X11 lock screens actually work, I realized the simplest way that'd suffice for me is to have the root window consume pointer and keyboard inputs via XGrabPointer and XGrabKeyboard respectively. Now I just swap DWM tags to a tag with no windows and press Mod4 and XK_l.
+
+## Formalization
+
+A proper ritual for accessing a system fulfils the following requirement:
+
+$r \in (A_i \cap B_i \cap ... \cup Z_i) - \Gamma_i$
+
+Where $A_i, B_i, ..., Z_i$ are the sets of individuals who should be able to perform the ritual to access the system, and $\Gamma_i$ is the collective intelligence of all individuals who shouldn't have access to the system.
+
+In the case of multi-user system's it's often useful to have one ritual per-person to simplify the process of creating a shared ritual, and for traceability.
+
+Notably, to fulfil the ritual, one mustn't be interrupted. Given this, time based rituals may be permissible in certain contexts. Things like proof of work, solving trivial problems, typing tests, and other things that are subject almost exclusively to time deltas can be rituals, assuming the amount of time it takes to perform the ritual is greater than the amount of time required to stop someone else from completing the ritual who isn't in the user set.
+
+
+## Uncommon Rituals
+
+Common rituals are passwords, biometrics, hardware keys, stuff like that. Uncommon ones, similar to what some use with their morning alarms, can perhaps be questions. If you are quite skillful at some set of problems, and are willing to solve them each time you are to unlock your system, these may be used in lieu of passwords, as long as they are rather complex.
+
+Another consideration is time. If I lock my screen, I'll generally be very close to my system while it's locked because I don't trust any lock screens that aren't actually logging out my user account, which don't seem to really exist in a sensible way. Given this, for me it'd actually be sensible to just have a small set of problems to solve that take some non-trivial amount of time that's larger than the amount of time I'll be away from my system.
+
+Another layer of protection I have in the time cost domain is my keyboard, split, colemak.
+
+### Knowledge
+
+An interesting solution to this problem is using knowledge. There are things about me my wife knows, things about me specific family members of mine know, and specific pieces of information about me friends know. Given these considerations, if I wanted my system to be accessible to multiple people, the selection of a ritual could be done on the basis of the union or our knowledge, differenced from the knowledge of everyone outside of the group.
+
+### Optimization
+
+We may want to optimize for time. Such an optimization likely dictates biometrics are to be used as the ritual, and every person who accesses the system will have their own ritual. If we are optimizing for fun or improvement, or some other objective there are far more options.
+
+#### Fun
+
+- Video games
+
+#### Improvement
+
+- Typing tests
+- Competitive programming
+- Math
+- Logic
+- Vim tutor
diff --git a/python/search-engines/search.csv b/python/search-engines/search.csv
@@ -92,3 +92,6 @@ degoog vs searxng,ddg,1780944913.4894233,False,0,False,3,-1,1
code completion model intelligence,startpage,1780963578.2915564,False,0,False,0,-1,0
code completion model intelligence,brave_search,1780963578.2915564,False,0,False,3,-1,0
code completion model intelligence,ddg,1780963578.2915564,False,0,False,0,1,0
+lock screen history,startpage,1789509690.6684983,False,0,False,1,-1,0
+lock screen history,brave_search,1789509690.6684983,False,0,False,4,-1,0
+lock screen history,ddg,1789509690.6684983,False,0,False,2,-1,0