This is not entirely correct. If you care about latency, then it doesn't matter on which major fault your application gets paused. Disabling swap protects you from data pages being evicted, but code pages can still be paged out.
If you care about latency, mlock() your memory, do not disable swap. Swap is good and gives the kernel an equal opportunity to evict data and code pages.
For GC enabled languages swap is universally bad. Some gc-pauses are indistinguishable from a system crash. It's a side effect on not having tightly specified memory limits.
I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.
If you want you application to stay in memory, then make it explicitly with mlock()/mlockall().
Disabling swap will just moves pressere elsewhere: to code pages. And evicted code page is no better: full stall while kernel loads that page from disk.
Unfortunately that's somewhat expected - regardless of implementation, the sole fact that the GC needs to somehow walk the entire tree to trace still referenced sections makes swap highly impractical -- even if you had not had 40ms stop-the-world pauses the LRU cache used by swap would get thrown away at every GC.
I believe that's actually the same reason why Apple stopped using GC in their frameworks in favour of automatic reference counting.
This. Memory bloat bugs are relatively easy to fix in Go, but sometimes it’s an adventure to remove swap bloat. And I think it’s worth removing all swap bloat!
I started wondering if there could be swap-aware GC, like first make the required page swapped in (not that there's any obvious API for that...) and only then pause the world?
One can just add APIs to the Linux kernel, and this would be a pretty straightforward one. (though it might be a little more difficult to avoid a syscall here)
I'm sure an agent can work on this and get some numbers with a day's worth of tokens.
You could do lots of interesting things with sufficiently deep inter-layer integration.
For example, why not swap out not by LRU page but by dense node clusters on the heap graph, maintaining in-memory summaries of inbound and outbound edges for liveness? If you do this, you don't have to swap the cluster in to do a GC involving it.
If the whole cluster becomes unreachable, you wouldn't even have to swap it back in to get rid of it: you'd just drop the swap reference and deem the swap space free.
I don't see anything this deeply integrated happening near-term, but it's fun to think about.
A GC latency SLO should include operating-system memory pressure. Otherwise, a page-fault problem will look like a collector problem and lead to the wrong fix.
The nasty bit is that swap doesn't just make the allocation slower; if GC metadata gets paged out, you have turned memory pressure into a stop-the-world latency spike.
In general when you tune knobs for GC, you pay for benefits in one area with sacrifices in another. Two big knobs to turn are pause latency and throughput. You probably wouldn’t want to go full “optimize for latency” because you’d end up with poor throughput. Also vice versa. Java’s reputation for poor GC performance is partly due to historical defaults that tune it for throughput.
Go’s GC is already a “concurrent mark-sweep garbage collector” and already has “extremely low mutator pause times, on the order of tens of microseconds”. It sounds like on-the-fly is just a different flavor of what Go already has.
It's a well known algorithm. Folks who do GCs for a living know about it. The folks who work on Go are surely aware of it. I'm assuming that they do not use it for a good reason, hence my question!
Fil-C's GC (Fil's Unbelievable Garbage Collector) uses an alternative on-the-fly algorithm, which I call Phil's Concurrent Marking.
Phil's Concurrent Marking differs from DLG in that it only requires a Djikstra barrier and uses a permagrey stack (something that Go used to do).
However, FUGC does clever things for coroutines (as in ucontexts, which Fil-C supports) - they are not permagrey; they only become grey if they execute. That's relevant to Go because Go moved away from permagrey stacks because of coroutine scan overheads, which the FUGC coroutine strategy might avoid.
But even if Go could not go back to permagrey, then the answer would be to use DLG, which would involve using the combined Yuasa+Dijstra barrier, which Go uses today anyway
ref counting is expensive in multi-threaded applications. Overall it would have worse performance. When it comes to predictability: deallocating a linked list (for instance) would have to deallocate all of the elements. Dealing with reference cycles is also not simple, either.
Discord learned this back in 2020 and published a great blog post on it: https://discord.com/blog/why-discord-is-switching-from-go-to...
That’s a great read. And an incredibly annoying “floating footer”
"It hurts when I do this"
"Stop doing that"
If you care about latency, disable swap. System wide or for the specific the cgroup.
This is not entirely correct. If you care about latency, then it doesn't matter on which major fault your application gets paused. Disabling swap protects you from data pages being evicted, but code pages can still be paged out.
If you care about latency, mlock() your memory, do not disable swap. Swap is good and gives the kernel an equal opportunity to evict data and code pages.
For GC enabled languages swap is universally bad. Some gc-pauses are indistinguishable from a system crash. It's a side effect on not having tightly specified memory limits.
I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.
If you want you application to stay in memory, then make it explicitly with mlock()/mlockall().
Disabling swap will just moves pressere elsewhere: to code pages. And evicted code page is no better: full stall while kernel loads that page from disk.
Plus, it wouldn't necessarily be that hard to selectively mlock the pages accessed in STW critical regions.
Unfortunately that's somewhat expected - regardless of implementation, the sole fact that the GC needs to somehow walk the entire tree to trace still referenced sections makes swap highly impractical -- even if you had not had 40ms stop-the-world pauses the LRU cache used by swap would get thrown away at every GC.
I believe that's actually the same reason why Apple stopped using GC in their frameworks in favour of automatic reference counting.
Reference counting is a GC algorithm, and no this isn't expected, it depends pretty much on the implementation.
Many make the mistake to think there is only one way to do a GC.
==> https://gchandbook.org/contents.html
This. Memory bloat bugs are relatively easy to fix in Go, but sometimes it’s an adventure to remove swap bloat. And I think it’s worth removing all swap bloat!
I started wondering if there could be swap-aware GC, like first make the required page swapped in (not that there's any obvious API for that...) and only then pause the world?
One can just add APIs to the Linux kernel, and this would be a pretty straightforward one. (though it might be a little more difficult to avoid a syscall here)
I'm sure an agent can work on this and get some numbers with a day's worth of tokens.
The API already exists: mlock.
You don’t need an api. Just touch the page.
Aren't madvise and mincore the APIs you want?
You could do lots of interesting things with sufficiently deep inter-layer integration.
For example, why not swap out not by LRU page but by dense node clusters on the heap graph, maintaining in-memory summaries of inbound and outbound edges for liveness? If you do this, you don't have to swap the cluster in to do a GC involving it.
If the whole cluster becomes unreachable, you wouldn't even have to swap it back in to get rid of it: you'd just drop the swap reference and deem the swap space free.
I don't see anything this deeply integrated happening near-term, but it's fun to think about.
A GC latency SLO should include operating-system memory pressure. Otherwise, a page-fault problem will look like a collector problem and lead to the wrong fix.
The nasty bit is that swap doesn't just make the allocation slower; if GC metadata gets paged out, you have turned memory pressure into a stop-the-world latency spike.
I use Go where i'd use Node, Ruby or Python.
I use Rust where i need low latency.
So, uh, mlock those pages?
Is there a reason why Go isn’t using on the fly GC, where there’s no STW at all?
What I found when I searched for "on the fly" is this: https://github.com/mthom/on-the-fly-gc — it looks like it has pauses, from the docs.
In general when you tune knobs for GC, you pay for benefits in one area with sacrifices in another. Two big knobs to turn are pause latency and throughput. You probably wouldn’t want to go full “optimize for latency” because you’d end up with poor throughput. Also vice versa. Java’s reputation for poor GC performance is partly due to historical defaults that tune it for throughput.
Go’s GC is already a “concurrent mark-sweep garbage collector” and already has “extremely low mutator pause times, on the order of tens of microseconds”. It sounds like on-the-fly is just a different flavor of what Go already has.
The classic on the fly GC algorithm is DLG, hilariously published in two papers, because the first one had a bug. Here's the second paper: https://caml.inria.fr/pub/papers/doligez_gonthier-gc-popl94....
It's a well known algorithm. Folks who do GCs for a living know about it. The folks who work on Go are surely aware of it. I'm assuming that they do not use it for a good reason, hence my question!
Fil-C's GC (Fil's Unbelievable Garbage Collector) uses an alternative on-the-fly algorithm, which I call Phil's Concurrent Marking.
I've documented it here: https://fil-c.org/fugc
Here's the source: https://github.com/pizlonator/fil-c/blob/deluge/libpas/src/l...
Phil's Concurrent Marking differs from DLG in that it only requires a Djikstra barrier and uses a permagrey stack (something that Go used to do).
However, FUGC does clever things for coroutines (as in ucontexts, which Fil-C supports) - they are not permagrey; they only become grey if they execute. That's relevant to Go because Go moved away from permagrey stacks because of coroutine scan overheads, which the FUGC coroutine strategy might avoid.
But even if Go could not go back to permagrey, then the answer would be to use DLG, which would involve using the combined Yuasa+Dijstra barrier, which Go uses today anyway
I don't get why people do not prefer reference counting, it has more predictable runtime performance
(though of course a swap is a swap - but you can "trigger" it depending on your memory or file access pattern)
Overhead per allocation, if you care about that, and reference circles.
ref counting is expensive in multi-threaded applications. Overall it would have worse performance. When it comes to predictability: deallocating a linked list (for instance) would have to deallocate all of the elements. Dealing with reference cycles is also not simple, either.
> more predictable runtime performance
Not really, reference counting can cause a single object deallocation to trigger an arbitrarily long chain of deallocations.
Yet another reason why "no runtime" is such a huge advantage.
You're probably coding for a 'runtime' in the broad sense: libc and your kernel.
Microcontroller! The platform is the peripherals.
What is stopping the author to use latest version of Go and Linux Kernel ?