Note [Generating code for top-level string literal bindings]
As described in Note [Compilation plan for top-level string literals] in GHC.Core, the core-to-core optimizer can introduce top-level Addr# bindings to represent string literals. The creates two challenges for the bytecode compiler: (1) compiling the bindings themselves, and (2) compiling references to such bindings. Here is a summary on how we deal with them: 1. Top-level string literal bindings are separated from the rest of the module. Memory is not allocated until bytecode link-time, the bc_strs field of the CompiledByteCode result records [(Name, ByteString)] directly. 2. When we encounter a reference to a top-level string literal, we generate a PUSH_ADDR pseudo-instruction, which is assembled to a PUSH_UBX instruction with a BCONPtrAddr argument. 3. The loader accumulates string literal bindings from loaded bytecode in the addr_env field of the LinkerEnv. 4. The BCO linker resolves BCONPtrAddr references by searching both the addr_env (to find literals defined in bytecode) and the native symbol table (to find literals defined in native code). This strategy works alright, but it does have one significant problem: we never free the memory that we allocate for the top-level strings. In theory, we could explicitly free it when BCOs are unloaded, but this comes with its own complications; see #22400 for why. For now, we just accept the leak, but it would nice to find something better.
References 1
Referenced by 9
- GHC.StgToByteCode call site ×2
- Allocating string literals GHC.ByteCode.Asm
- GHC.ByteCode.Asm call site
- GHC.ByteCode.Instr call site
- GHC.ByteCode.Linker call site
- GHC.ByteCode.Types call site
- GHC.Linker.Loader call site
- GHC.Linker.Types call site