Repository navigation
Conversation
|
@luames please dont do two seperate txtar testfiles for this change |
7a116eb to
a523c75
Compare
|
@mvdan opinions on this? |
a523c75 to
6d7e7eb
Compare
mvdan
left a comment
There was a problem hiding this comment.
I see this doesn't reuse the algorithms from literal obfuscation. Why not? At the end of the day, a var string = "lots of bytes" is not different from the same amount of bytes via a //go:embed.
Assuming the answer is that it's because those algorithms don't scale so they have a size limit, then I think the answer is to offer simpler or cheaper algorithms that can work with much larger sizes, and then use them consistently for both literals and embeds.
Also, did you measure how this affects the performance of builds, both the build itself and the running of a binary?
In principle adding this feature to garble SGTM, but ideally we don't add one more "mode" flag and ideally we don't end up with two different mechanisms to obfuscate-but-not-encrypt strings, small or large.
|
@luames please address the feedback mvdan left for you |
|
Addressed both review comments in 6f88b44. Removed the Python benchmark script and formatted all three Go source entries in The non-short |
|
@luames please rebase |
6f88b44 to
820ed26
Compare
|
Rebased onto current The literal unit tests, non-short embed/literals/help scripts, and |
|
@luames please rebase |
Issue burrowers#194 has remained open because the compiler places embedded files directly in object data, outside the existing literal transformation. An opt-in compiler-source transformation now reads the build system embed configuration and replaces embedded string and byte variables with deterministic decoded initializers, so their original contents are absent from the resulting binary without changing the source files. Named string and byte-slice types retain their declared type and package initialization dependencies. Cache hashing accounts for the new flag. embed.FS remains unsupported under -embed and fails explicitly rather than emitting plaintext; large files can still produce expensive generated source. The follow-up FS and large-asset work remains tracked in burrowers#194. Refs burrowers#194
The initial burrowers#194 PR covered embedded string and byte variables but rejected embed.FS, leaving directory trees and callers using the standard fs.FS interface unsupported. Rewrite resolved -embedcfg file paths to encoded temporary assets and adapt the embed package at compile time so Open, ReadFile, ReadAt, Seek, Stat, ReadDir, and WalkDir still return original contents and sizes. Keep the compiler-known embed.file layout unchanged, type-check the patched package before identifier obfuscation, and assert the expected Go source layout to fail closed on toolchain drift. The integration test exercises nested directories, an embed.FS alias, a non-obfuscated dependency, and combination with -literals. Original assets remain untouched; the decoding key in the binary makes this obfuscation rather than encryption. Refs burrowers#194
The burrowers#194 change introduced separate scripts for embedded variables and embed.FS, forcing CI to run another independent Garble integration fixture. Exercise string, byte, and filesystem embedding together in one script while retaining the named-type, alias, initialization, non-obfuscated dependency, GOGARBLE, -literals, runtime output, and binary marker checks. Refs burrowers#194
Embedded data for burrowers#194 previously used a separate decoder and -embed mode, while ordinary literals above 2 KiB remained plaintext because per-byte AST algorithms did not scale. Follow the review by extending -literals to embeds and sharing both small-value algorithms and a linear-time blob algorithm for large values. Resolve embedded strings and bytes into the literal pass, retaining named types and initialization dependencies. Use the same blob encoder and decoder for embed.FS assets, sorting filenames before consuming randomness. Constrain decoded slice capacity to its length so generated byte literals preserve their original behavior. Consolidate all embed coverage into embed.txtar and add empty/tiny data, large binary fuzz seeds, and bounded-AST regression coverage. Document that keys remain recoverable and filesystem decoding repeats per open. Add a benchmark that verifies content hashes and records rebuild/cache CPU and wall time, initialization, reads, allocations, and binary size. For a 1 MiB payload on a shared Linux/amd64 container, warm-dependency rebuild CPU seconds for Garble versus -literals were 1.90/2.12 for a literal, 1.74/2.01 for an embedded string, 1.76/2.13 for embedded bytes, and 1.83/1.84 for embed.FS. Single rebuild samples are not statistically significant. Five-run median initialization was 3.53, 5.22, and 8.13 ms for decoded literal/string/bytes. FS reads averaged 19.00 versus 54.96 ms and allocated about 1 versus 3 MiB per read. Each binary grew 4096 bytes. Validated with Go 1.27.0: internal/literals tests, non-short embed, literals and help script tests, go vet ./..., git diff --check, and all 12 benchmark variants with content and plaintext-marker assertions. Refs burrowers#194
Keep the embed contribution within the project tooling conventions by removing the standalone Python benchmark. The performance measurements already recorded in the PR remain historical evidence, not a dependency for building or testing Garble. Format all three Go source entries in embed.txtar with go/format without changing its assets, expected output, or test commands. The non-short TestScript/embed regression passes with Go 1.27.0. Refs burrowers#194.
820ed26 to
96fb85a
Compare
|
Rebased onto current Literal unit tests, the non-short embed/literals/help scripts, and |
Refs #194. The compiler stores
//go:embedcontents outside Garble's ordinary literal pass. This change brings embeddedstring,[]byte, andembed.FScontents under-literals; there is no separate embed flag.Embedded string and byte variables now go through
internal/literals, using its existing algorithms for small values. A shared blob algorithm handles values above 2 KiB with linear-time encoding and decoding and a bounded number of AST nodes. Ordinary large literals use it too, rather than remaining plaintext. Theembed.FSoverlay uses the same encoder and decoder, keeping assets in compiler data and decoding them when opened. Named types, initialization dependencies, filesystem metadata, seek, andReadAtbehavior remain covered. The branch also incorporates current master, including removal of GOGARBLE selection.This is obfuscation, not encryption. Keys and decoders remain in the binary, and embedded filenames remain visible. Large variable initializers still contain an encoded Go string. Embedded variables decode during initialization; filesystem opens decode again and allocate temporary buffers.
The existing
embed.txtarfixture now contains all embed coverage; the separate obfuscation fixture is removed. Tests exercise normal and-literalsbuilds, named types, aliases, initialization dependencies, empty and tiny contents, nested files, cross-package filesystem assets, runtime output, and plaintext-marker absence. Large-literal unit and fuzz seeds cover binary data and slice length/capacity preservation.Verification
With Go 1.27.0 on Linux/amd64:
go test ./internal/literalspassed.go test -timeout=30m -v -run '^TestScript/(embed|literals|help)$' -count=1 .passed, including the non-short reproducibility assertions.go vet ./...andgit diff --checkagainst current master passed.CI for
d7cd406passed all six jobs, including Linux, macOS, Windows, 386, race, and third-party checks.The 1 MiB performance sample below was collected on
d7cd406for literal, embedded string, embedded bytes, and filesystem variants under plain Go, Garble, and Garble with-literals. Every variant verified the decoded SHA-256; every obfuscated binary rejected the plaintext marker. The one-off Python measurement script was removed in response to review and is not part of this contribution.CI for
6f88b44passed all six jobs.After formatting every Go source entry in
embed.txtar,go test -timeout=30m -v -run '^TestScript/embed$' -count=1 .passed with Go 1.27.0 on6f88b44.After rebasing onto
2d6ef16, the literal unit tests, non-short embed/literals/help scripts,go vet ./..., andgit diff --checkpassed on96fb85awith Go 1.27.0. A multibyte literal above 2 KiB also preserved content, length, and exact capacity for plain and named byte/rune slices compared with plain Go. CI for this head passed all six jobs.Performance sample
The benchmark records warm-dependency application rebuilds, fully cached builds, child CPU time, binary size, initialization, process execution, and reads. These are single rebuild samples from a shared, CPU-limited container, not statistically significant comparisons. Runtime figures are medians of five process runs. The payload is 1 MiB in each variant, and dependencies were warmed before measurement.
-literals-literals-literalsThese measurements use the merged
d7cd406tree. The comparison enables-literalsthroughout the application and standard-library dependencies, so it includes ordinary literal obfuscation as well as embedding. Binary sizes grew from 2.85 to 5.37 MiB for the literal/string/bytes variants and from 2.95 to 5.53 MiB for the filesystem variant; that is not the isolated cost of embedding.Read timings average 100 reads per process. The string cases copy to bytes; the byte case reuses its initialized slice. Filesystem reads allocate about 1 MiB without literal obfuscation and 3 MiB with it, reflecting the temporary decode buffers. The embedded-string initialization sample was especially noisy and should not be treated as a stable latency estimate. The recorded wall-clock build measurements were too noisy here to support a speed claim.
Implemented by Hermes Agent using gpt-6.1-sol, acting on behalf of @luantak.