diff --git a/.gitignore b/.gitignore index 9d60e124..04a2f291 100644 --- a/.gitignore +++ b/.gitignore @@ -2,4 +2,5 @@ docs/.hugo_build.lock docs/public/ *.swp /cmd/write-test-data/write-test-data +/write-test-data .lycheecache diff --git a/MaxMind-DB-spec.md b/MaxMind-DB-spec.md index 37a854e1..38b68c04 100644 --- a/MaxMind-DB-spec.md +++ b/MaxMind-DB-spec.md @@ -310,6 +310,10 @@ A pointer to another part of the data section's address space. The pointer will point to the beginning of a field. It is illegal for a pointer to point to another pointer. +Because several pointers can share one target, following pointers naively can +decode far more data than a record's size suggests. See +[Reader Resource Limits](#reader-resource-limits). + Pointer values start from the beginning of the data section, _not_ the beginning of the file. Pointers in the metadata start from the beginning of the metadata section. @@ -537,6 +541,137 @@ are ignored. This means that we are limited to 4GB of address space for pointers, so the data section size for the database is limited to 4GB. +## Reader Resource Limits + +A crafted or corrupt data section can make a reader use far more time and memory +than a data entry's encoded size suggests. A value can be a pointer, and several +pointers can share one target. A small data section can therefore describe a +structure that is huge, or effectively infinite, when fully expanded. + +A reader should apply resource controls to operations that decode +attacker-controlled values. The controls should bound the reader's time and +memory before it allocates, traverses, or copies an attacker-chosen amount of +data. The examples in this section are implementation guidance. They do not +define whether an encoding conforms to the file format. + +A caller-requested decode operation can include a lookup, path selection and +decoding of the selected value, or metadata decoding while opening a database. A +reader that validates a whole database can apply a bound to each data entry if +the work outside those entries is also bounded. + +A reader can apply the depth, value-count, and payload limits below. It can +instead combine them or use another strategy that bounds the same risks, such +as: + +- Safe memoization of pointer targets, including cycle handling. +- Schema-directed decoding into a destination that is finite and nonrecursive, + and that skips unknown values without following pointers. +- A weighted work budget. + +A reader that delegates decoding to caller-defined code should document where +that responsibility transfers. + +A finite, nonrecursive destination can provide an inherent bound when its +recognized fields have bounded shapes and unknown values are skipped without +following their pointers. A specific typed destination does not provide that +bound by itself. It can still contain recursive types, attacker-sized +collections, repeated recognized fields, or dynamically shaped values. A +schema-aware reader can apply tighter semantic limits where it knows a +collection's valid size is small. + +### Maximum Data Structure Depth + +A reader that uses a nesting counter should limit how deeply it decodes one data +entry. The depth increases by one each time the reader enters a map or an array, +or follows a pointer. If the depth exceeds 512, the reader should stop and +reject the data entry. + +This limit stops unbounded recursion, including a pointer cycle and a structure +nested deeper than any real database needs, and it is a portable default. +Iterative decoding avoids call-stack exhaustion, but it must still detect cycles +and bound nesting and traversal work. + +### Maximum Decoded Value Count + +The depth limit does not bound the total amount of work. Arrays and maps can +contain many pointers to the same target. Unless the reader safely reuses the +target, it decodes that target once for every pointer occurrence. In the binary +fan-out example used by the test data, each array contains two pointers to the +level below, so each lower level doubles the work. Decoding the top value takes +exponentially more operations than the depth suggests. A data entry under one +kilobyte can therefore take longer to decode than any real workload allows. + +A reader can bound this by counting the values it decodes and stopping if the +count exceeds a fixed limit. One flat accounting rule is: + +- Charge every logical value occurrence once. The root is one occurrence. An + array or map is one occurrence in addition to its children. Map keys and map + values are separate occurrences. +- Do not charge a pointer separately from its resolved value. Charge the value + at the position where the pointer occurs. Under this rule, a value resolved + from a memoized target is still charged once for each logical occurrence. + +Readers with structural or weighted bounds may account differently. + +If a destination uses a container's declared length to allocate storage, +counting only after that allocation may be too late. A reader can reserve the +declared children against its value budget first, cap the allocation separately, +grow storage incrementally, or use another approach with an equivalent bound. + +The bound should cover the whole operation the caller requested. It should not +restart between internal phases. For example, path navigation and decoding the +selected value can share one budget or use another end-to-end control. Bounding +each navigation step on its own does not bound their cumulative work. + +A limit of 65,536 (2\*\*16) values is recommended. The largest data entries +MaxMind produces decode a few hundred values, so this leaves a wide margin. + +### Bounding Expanded Payload + +The value-count strategy limits how many values the reader decodes from a data +entry, not how many bytes those values contain. A single string or bytes value +can be up to 16,843,036 bytes. An array of 65,535 pointers to one such value +stays at the recommended value limit under the flat rule above. It can still +describe more than 1 TiB of repeated data in a file barely larger than the value +itself. + +A reader that copies, validates, or allocates these values should bound the +total string and bytes data it materializes for one caller-requested decode +operation. The right method and limit depend on the reader's language and API, +so this specification does not require a single limit. As guidance, the largest +data entries MaxMind produces hold about a kilobyte of such data, so a few +megabytes is generous. + +Borrowing bytes from the data section or safely memoizing pointer targets can +bound the reader's own materialization. The decoded result may still contain +many logical references to the same value. A binding, serializer, or conversion +to owning values that copies each occurrence should bound that downstream work. + +Payload that a reader skips without copying, validating, or allocating need not +consume a materialization budget, as long as the traversal to skip it is bounded +separately. A reader selecting part of a data entry likewise need not expand an +unrequested pointer target. A reader whose API or validation rules require that +work should bound it. + +A reader can bound this in several ways: + +- Charge each value's size against a per-operation budget wherever it is + decoded, including data stored inline in a container that a pointer targets. + Stop when the total exceeds a limit. Re-decoding a pointer target charges its + data again, which is what bounds the amplification. +- Fold these concerns into a unified weighted budget. It can account for + container entries traversed, pointer expansion when the target is not safely + reused, map-key and selector work, and strings and bytes materialized. A small + fixed-width scalar can cost little or nothing once the work needed to reach it + is bounded. +- Safely memoize decoded pointer targets, handling cycles and in-progress + targets, so a shared target is materialized once. + +When a reader's resource controls reject a decode, it should return an error +rather than a partial result. It may treat a value outside its documented +resource profile as invalid, and may make limits configurable for unusually +large valid data entries. + ## Reference Implementations ### Writer diff --git a/README.md b/README.md index d5d20bdf..43616d01 100644 --- a/README.md +++ b/README.md @@ -3,6 +3,12 @@ subnets (IPv4 or IPv6). This repository contains the spec for that format as well as test databases. +Some structurally well-formed databases under `test-data/` are deliberately +hostile and can exhaust an unprotected reader. Do not fully decode every fixture +without resource controls. See the +[denial-of-service test data](test-data/README.md#denial-of-service-test-data) +documentation. + # Generating Test Data The `write-test-data` command generates the MMDB test files under `test-data/` diff --git a/cmd/write-test-data/main.go b/cmd/write-test-data/main.go index 638deea0..d7a7479c 100644 --- a/cmd/write-test-data/main.go +++ b/cmd/write-test-data/main.go @@ -82,6 +82,11 @@ func main() { os.Exit(1) } + if err := w.WritePointerDecoderDoSTestDB(); err != nil { + fmt.Printf("writing pointer decoder DoS test databases: %+v\n", err) + os.Exit(1) + } + if err := w.WriteGeoIP2TestDB(); err != nil { fmt.Printf("writing GeoIP2 test databases: %+v\n", err) os.Exit(1) diff --git a/go.mod b/go.mod index ce3e46d2..885bb7bb 100644 --- a/go.mod +++ b/go.mod @@ -4,10 +4,8 @@ go 1.25.0 require ( github.com/maxmind/mmdbwriter v1.2.0 + github.com/oschwald/maxminddb-golang/v2 v2.1.1 go4.org/netipx v0.0.0-20260823151212-3075585bcbeb ) -require ( - github.com/oschwald/maxminddb-golang/v2 v2.1.1 // indirect - golang.org/x/sys v0.38.0 // indirect -) +require golang.org/x/sys v0.38.0 // indirect diff --git a/pkg/writer/pointerdos.go b/pkg/writer/pointerdos.go new file mode 100644 index 00000000..8069d5fd --- /dev/null +++ b/pkg/writer/pointerdos.go @@ -0,0 +1,502 @@ +package writer + +import ( + "fmt" + "os" + "path/filepath" +) + +// pointerDoSDepth is the nesting depth of the fan-out structure. Each level +// adds a factor of two to the number of leaf decodes an unprotected reader that +// re-decodes every target per referencing path performs. The 2**40 leaf +// decodes, and Θ(2**40) total decode operations, make a sub-kilobyte file +// impossible to decode that way. Memoization is one defense. A cumulative work +// budget or a schema-directed path also avoids the blow-up. +const pointerDoSDepth = 40 + +const ( + // pointerDoSBuildEpoch is fixed so the generated files are reproducible. + pointerDoSBuildEpoch = 1_000_000_000 + + pointerDoSFixtureFilename = "MaxMind-DB-test-pointer-decoder-dos.mmdb" + pointerDoSIPv6FixtureFilename = "MaxMind-DB-test-pointer-decoder-dos-ipv6.mmdb" + pointerValueLimitFixtureFilename = "MaxMind-DB-test-decoder-value-limit-pointer-heavy.mmdb" + payloadDoSFixtureFilename = "MaxMind-DB-test-payload-amplification-dos.mmdb" + worstCasePayloadFixtureFilename = "MaxMind-DB-test-payload-amplification-dos-worst-case.mmdb" + stringPayloadFixtureFilename = "MaxMind-DB-test-payload-amplification-dos-string.mmdb" + valueLimitFixtureFilename = "MaxMind-DB-test-decoder-value-limit.mmdb" + valueLimitOverFixtureFilename = "MaxMind-DB-test-decoder-value-limit-over.mmdb" + payloadLimitFixtureFilename = "MaxMind-DB-test-decoder-payload-limit.mmdb" + payloadLimitOverFixtureFilename = "MaxMind-DB-test-decoder-payload-limit-over.mmdb" + metadataLimitFixtureFilename = "MaxMind-DB-test-metadata-payload-limit.mmdb" + decodePathBudgetFixtureFilename = "MaxMind-DB-test-decode-path-shared-budget.mmdb" +) + +// writeDataPointer writes a data-section pointer with a one-byte payload, +// which addresses offsets 0..2047. The value is relative to the start of the +// data section, which is where the reader sets the pointer base. +func writeDataPointer(target int) []byte { + if target < 0 || target > 2047 { + panic(fmt.Sprintf( + "data pointer target %d is out of range for a one-byte payload (0..2047)", + target, + )) + } + return []byte{ + (1 << 5) | byte((target>>8)&0x7), + byte(target & 0xFF), + } +} + +// dataRecordValue24 converts a data-section offset to a 24-bit search-tree +// record value. A search-tree record value is biased by the node count and the +// data-section separator, so it is not the bare data-section offset that a +// data-section pointer uses. This adds that bias. +func dataRecordValue24(nodeCount uint32, dataOffset int) uint32 { + if dataOffset < 0 { + panic(fmt.Sprintf("negative data-section offset: %d", dataOffset)) + } + recordValue := uint64(nodeCount) + dataSeparatorSize + uint64(dataOffset) + if recordValue > maximum24BitSearchTreeValue { + panic(fmt.Sprintf( + "data record value %d does not fit in a 24-bit search-tree record", + recordValue, + )) + } + return uint32(recordValue) +} + +// buildPointerFanOutData builds the data section: nested arrays, each holding +// two pointers to the node below. It is laid out leaf first, so the leaf is at +// offset 0 and each array points back to the level below it. A decoder that +// re-decodes a shared pointer target once per referencing path performs +// 2**depth leaf decodes and Θ(2**depth) total decode operations from these few +// hundred bytes; a decoder that memoizes resolved targets decodes it in linear +// time. mmdbwriter cannot produce this shape because its +// own deduplication fans out while computing the file, so the bytes are written +// directly. top is the data-section offset of the outermost array. +func buildPointerFanOutData(depth int) (data []byte, top int) { + data = []byte{0xA0} // uint16 with value 0 + prev := 0 + for range depth { + offset := len(data) + data = append(data, 0x02, 0x04) // array (extended type 11), size 2 + data = append(data, writeDataPointer(prev)...) + data = append(data, writeDataPointer(prev)...) + prev = offset + } + return data, prev +} + +const ( + recommendedValueLimit = 1 << 16 + fixturePayloadLimit = 1 << 21 + decodePathDecoyCount = 511 + decodePathSharedKeySize = 4096 + decodePathValueSize = 4091 + + // payloadScalarSize is the size of the single shared bytes value that the + // pointers target. Size code 30 covers 285..65820 bytes. + payloadScalarSize = 1<<16 - 1 + // payloadPointerCount pointers each re-materialize the shared value, so a + // reader that copies the target for every pointer allocates + // payloadPointerCount*payloadScalarSize bytes, about 512 MiB, from a file of + // a few tens of kilobytes. + payloadPointerCount = 8192 + // worstCasePointerCount is the largest fan-out a reader whose only defense is + // the value count still accepts. Under flat value accounting the array plus + // its elements decode to worstCasePointerCount+1 = 65,536 values, which meets + // the recommended limit without exceeding it, so that reader does not reject + // it. Copying each target materializes worstCasePointerCount*payloadScalarSize + // bytes, about 4 GiB, from a file of about 192 KiB. A limit on the copied + // bytes stops this, as can safe reuse or structural rejection. + worstCasePointerCount = 65535 +) + +// MMDB scalar type codes used by the payload fixtures: 2 is a UTF-8 string and +// 4 is bytes. +const ( + scalarTypeString byte = 2 + scalarTypeBytes byte = 4 +) + +func writeScalar(buf []byte, size int, scalarType byte) int { + if size < 0 || size > maximumDataStructureSize { + panic(fmt.Sprintf( + "scalar size %d is outside the supported range 0..%d", + size, + maximumDataStructureSize, + )) + } + if scalarType != scalarTypeString && scalarType != scalarTypeBytes { + panic(fmt.Sprintf("unsupported scalar type %d", scalarType)) + } + + pos := 1 + switch { + case size <= 28: + buf[0] = (scalarType << 5) | byte(size&0x1F) + case size <= 284: + buf[0] = (scalarType << 5) | 29 + buf[1] = byte(size - 29) + pos++ + case size <= maximumSizeCode30: + buf[0] = (scalarType << 5) | 30 + encoded := size - 285 + buf[1] = byte((encoded >> 8) & 0xFF) + buf[2] = byte(encoded & 0xFF) + pos += 2 + default: + buf[0] = (scalarType << 5) | 31 + encoded := size - minimumSizeCode31 + buf[1] = byte((encoded >> 16) & 0xFF) + buf[2] = byte((encoded >> 8) & 0xFF) + buf[3] = byte(encoded & 0xFF) + pos += 3 + } + clear(buf[pos : pos+size]) + return pos + size +} + +func writeArrayHeader(buf []byte, size int) int { + if size < 0 || size > maximumDataStructureSize { + panic(fmt.Sprintf( + "array size %d is outside the supported range 0..%d", + size, + maximumDataStructureSize, + )) + } + switch { + case size <= 28: + buf[0] = byte(size & 0x1F) + buf[1] = 4 + return 2 + case size <= 284: + buf[0] = 29 + buf[1] = 4 + buf[2] = byte(size - 29) + return 3 + case size <= maximumSizeCode30: + buf[0] = 30 + buf[1] = 4 + encoded := size - 285 + buf[2] = byte((encoded >> 8) & 0xFF) + buf[3] = byte(encoded & 0xFF) + return 4 + default: + return writeLargeArray(buf, uint32(size)) + } +} + +// buildPayloadAmplificationData builds a data section holding one large scalar +// value (string or bytes, per scalarType) at offset 0 followed by an array of +// pointers that all target it. The value count and depth limits do not bound +// this: the array stays well under the value limit, yet a reader that copies +// each pointer's target materializes payloadPointerCount*payloadScalarSize +// bytes. top is the offset of the array. +func buildPayloadAmplificationData( + pointerCount int, + scalarType byte, +) (data []byte, top int) { + data = make([]byte, payloadScalarSize+16+pointerCount*2) + pos := writeScalar(data, payloadScalarSize, scalarType) + // The payload must not be left as NUL. A NUL-filled string measures length 0 + // through strlen, so a binding that copies C strings would skip the copy path + // this fixture exists to exercise. + for i := pos - payloadScalarSize; i < pos; i++ { + data[i] = 'a' + } + + top = pos + pos += writeArrayHeader(data[pos:], pointerCount) + for range pointerCount { + pos += copy(data[pos:], writeDataPointer(0)) + } + return data[:pos], top +} + +// buildValueLimitData produces one array node followed by pointerCount scalar +// nodes. The scalar has no string/bytes payload, so only the decoded-value +// budget determines whether the record is accepted. +func buildValueLimitData(pointerCount int) (data []byte, top int) { + data = make([]byte, 16+pointerCount*2) + data[0] = 0xA0 // uint16 with value 0 + pos := 1 + top = pos + pos += writeArrayHeader(data[pos:], pointerCount) + for range pointerCount { + pos += copy(data[pos:], writeDataPointer(0)) + } + return data[:pos], top +} + +// buildPayloadLimitData produces 32 references to a 65,535-byte value and one +// reference to a 32- or 33-byte value. Those totals are exactly 2 MiB and one +// byte over 2 MiB respectively, while the decoded-value count remains tiny. +func buildPayloadLimitData(smallSize int) (data []byte, top int) { + const largePointerCount = 32 + data = make([]byte, payloadScalarSize+smallSize+128) + pos := writeScalar(data, smallSize, scalarTypeBytes) + largeOffset := pos + pos += writeScalar(data[pos:], payloadScalarSize, scalarTypeBytes) + top = pos + pos += writeArrayHeader(data[pos:], largePointerCount+1) + for range largePointerCount { + pos += copy(data[pos:], writeDataPointer(largeOffset)) + } + pos += copy(data[pos:], writeDataPointer(0)) + return data[:pos], top +} + +// buildDecodePathSharedBudgetData produces a map whose first 511 keys point to +// one shared 4 KiB string. Reading those keys and the final inline "target" key +// consumes 2,093,062 bytes. The selected 4,091-byte value then takes the total +// to 2,097,153 bytes, exactly one byte above the 2 MiB example payload budget +// used by these fixtures. This verifies that path navigation and selected-value +// decoding share one budget. +func buildDecodePathSharedBudgetData() (data []byte, top int) { + data = make([]byte, 16*1024) + pos := writeScalar(data, decodePathSharedKeySize, scalarTypeString) + for i := pos - decodePathSharedKeySize; i < pos; i++ { + data[i] = 'k' + } + + top = pos + pos += writeMap(data[pos:], decodePathDecoyCount+1) + for range decodePathDecoyCount { + pos += copy(data[pos:], writeDataPointer(0)) + data[pos] = 0 + data[pos+1] = 7 // extended type 14, size 0: false + pos += 2 + } + + const target = "target" + pos += writeScalar(data[pos:], len(target), scalarTypeString) + copy(data[pos-len(target):pos], target) + pos += writeScalar(data[pos:], decodePathValueSize, scalarTypeString) + for i := pos - decodePathValueSize; i < pos; i++ { + data[i] = 'v' + } + + return data[:pos], top +} + +// buildPayloadAmplificationDB builds a minimal valid IPv4 MMDB whose single +// search-tree node resolves every supported lookup to the payload +// amplification record. See buildPayloadAmplificationData. +func buildPayloadAmplificationDB(pointerCount int, scalarType byte) []byte { + data, top := buildPayloadAmplificationData(pointerCount, scalarType) + return buildSingleRecordDB(data, top) +} + +func buildSingleRecordDB(data []byte, top int) []byte { + const nodeCount = 1 + // A data record value is the data-section offset plus the node count plus the + // 16-byte data section separator, which is how a reader recovers the offset: + // offset = recordValue - nodeCount - dataSeparatorSize. + recordValue := dataRecordValue24(nodeCount, top) + + buf := make([]byte, 1024+len(data)) + pos := 0 + pos += writeSearchTree(buf[pos:], recordValue) + pos += dataSeparatorSize + pos += copy(buf[pos:], data) + pos += writeMetadataBlock(buf[pos:], nodeCount, pointerDoSBuildEpoch) + return buf[:pos] +} + +func buildValueLimitDB(pointerCount int) []byte { + data, top := buildValueLimitData(pointerCount) + return buildSingleRecordDB(data, top) +} + +func buildPayloadLimitDB(smallSize int) []byte { + data, top := buildPayloadLimitData(smallSize) + return buildSingleRecordDB(data, top) +} + +func buildDecodePathSharedBudgetDB() []byte { + data, top := buildDecodePathSharedBudgetData() + return buildSingleRecordDB(data, top) +} + +func writeMetadataLimitBlock(buf []byte, nodeCount uint32, buildEpoch uint64) int { + pos := 0 + copy(buf[pos:], metadataMarker) + pos += len(metadataMarker) + metadataStart := pos + pos += writeMap(buf[pos:], len(metadataKeysStandard)) + + var databaseTypeOffset int + for _, key := range metadataKeysStandard { + pos += writeMetaKey(buf[pos:], key) + switch key { + case "binary_format_major_version": + pos += writeUint16(buf[pos:], 2) + case "binary_format_minor_version": + pos += writeUint16(buf[pos:], 0) + case "build_epoch": + pos += writeUint64(buf[pos:], buildEpoch) + case "database_type": + databaseTypeOffset = pos - metadataStart + pos += writeScalar(buf[pos:], payloadScalarSize, scalarTypeString) + for i := pos - payloadScalarSize; i < pos; i++ { + buf[i] = 'T' + } + case "description": + pos += writeMap(buf[pos:], 0) + case "ip_version": + pos += writeUint16(buf[pos:], 4) + case "languages": + // The pointers below target the database_type string, so that key must + // already be written. Offset 0 is the metadata map's own control byte, + // which would silently make this a pointer cycle instead of a payload + // fixture. + if databaseTypeOffset == 0 { + panic("languages must be written after database_type") + } + const pointerCount = 33 + pos += writeArrayHeader(buf[pos:], pointerCount) + for range pointerCount { + pos += copy(buf[pos:], writeDataPointer(databaseTypeOffset)) + } + case "node_count": + pos += writeUint32(buf[pos:], nodeCount) + case "record_size": + pos += writeUint16(buf[pos:], 24) + default: + panic("unknown metadata key: " + key) + } + } + return pos +} + +func buildMetadataLimitDB() []byte { + const nodeCount = 1 + const recordValue = nodeCount + dataSeparatorSize + buf := make([]byte, 128*1024) + pos := writeSearchTree(buf, recordValue) + pos += dataSeparatorSize + pos += writeMap(buf[pos:], 0) + pos += writeMetadataLimitBlock(buf[pos:], nodeCount, pointerDoSBuildEpoch) + return buf[:pos] +} + +// buildPointerFanOutDB builds a minimal valid IPv4 MMDB whose single search-tree +// node resolves every supported lookup to the fan-out data section. See +// buildPointerFanOutData. +func buildPointerFanOutDB(depth int) []byte { + data, top := buildPointerFanOutData(depth) + return buildSingleRecordDB(data, top) +} + +// buildPointerFanOutAllSpaceDB builds a conventional IPv6 MMDB that maps the +// entire address space to the fan-out data record, so opening it, looking up +// any address, and decoding the result reproduces the denial of service. A +// lookup only finds the record; decoding it is what expands the fan-out. This +// is the form a third-party implementation is most likely to test, since +// implementations commonly test with database files rather than bare data +// sections. +// +// The search tree is a spine down the all-zeros path. At every node the +// one-branch resolves to the data record and the zero-branch descends to the +// next node, and the final node resolves both branches to the record. The spine +// is 96 nodes deep so that an IPv4 lookup, which follows the ::0/96 prefix, +// reaches the final node and the record too. +func buildPointerFanOutAllSpaceDB(depth int) []byte { + data, top := buildPointerFanOutData(depth) + + const ipv4PrefixBits = 96 + const nodeCount = ipv4PrefixBits + 1 // nodes 0..96 + badRecord := dataRecordValue24(nodeCount, top) + + const recordPairSize = 6 // two 24-bit records per node + buf := make([]byte, nodeCount*recordPairSize+dataSeparatorSize+len(data)+256) + pos := 0 + for i := range uint32(nodeCount) { + leftRecord := badRecord + if i < ipv4PrefixBits { + leftRecord = i + 1 // descend the all-zeros path + } + pos += writeSearchTreeRecords(buf[pos:], leftRecord, badRecord) + } + pos += dataSeparatorSize + pos += copy(buf[pos:], data) + pos += writeMetadataBlockWithKeyOrder( + buf[pos:], + nodeCount, + pointerDoSBuildEpoch, + 6, + metadataKeysStandard, + ) + return buf[:pos] +} + +// pointerValueLimitDepth is the deepest fan-out whose flat value count a reader +// at the recommended limit still accepts. A binary fan-out of depth d decodes +// 2**(d+1) - 1 values, so depth 15 gives 65,535, one below the limit, and depth +// 16 would give 131,071. Depth 15 is therefore the only depth that lands on the +// boundary from below. +const pointerValueLimitDepth = 15 + +// pointerDoSFixtures returns every fixture this file generates, keyed by +// filename. The writer and the byte-equality test share it so a new fixture +// cannot be added to one without the other. +func pointerDoSFixtures() map[string][]byte { + return map[string][]byte{ + pointerDoSFixtureFilename: buildPointerFanOutDB(pointerDoSDepth), + pointerDoSIPv6FixtureFilename: buildPointerFanOutAllSpaceDB( + pointerDoSDepth, + ), + pointerValueLimitFixtureFilename: buildPointerFanOutDB( + pointerValueLimitDepth, + ), + payloadDoSFixtureFilename: buildPayloadAmplificationDB( + payloadPointerCount, + scalarTypeBytes, + ), + worstCasePayloadFixtureFilename: buildPayloadAmplificationDB( + worstCasePointerCount, + scalarTypeBytes, + ), + stringPayloadFixtureFilename: buildPayloadAmplificationDB( + payloadPointerCount, + scalarTypeString, + ), + valueLimitFixtureFilename: buildValueLimitDB( + recommendedValueLimit - 1, + ), + valueLimitOverFixtureFilename: buildValueLimitDB( + recommendedValueLimit, + ), + payloadLimitFixtureFilename: buildPayloadLimitDB( + fixturePayloadLimit - 32*payloadScalarSize, + ), + payloadLimitOverFixtureFilename: buildPayloadLimitDB( + fixturePayloadLimit - 32*payloadScalarSize + 1, + ), + metadataLimitFixtureFilename: buildMetadataLimitDB(), + decodePathBudgetFixtureFilename: buildDecodePathSharedBudgetDB(), + } +} + +// WritePointerDecoderDoSTestDB writes the databases that exercise the +// data-section pointer denial of service. Two exercise the fan-out: a minimal +// single-node database and a conventional IPv6 database that maps all of the +// address space to the fan-out record. Three exercise payload amplification: a +// moderate and a worst-case fixture with a shared bytes value, and one with a +// shared string value, the type most bindings copy into a native string. It +// also writes exact boundary fixtures for the recommended value-count limit and +// the 2 MiB example payload budget, plus a metadata fixture that exceeds that +// budget while opening. A small path fixture verifies that navigation and +// selected-value decoding share the same budget. See buildPointerFanOutData and +// buildPayloadAmplificationData. +func (w *Writer) WritePointerDecoderDoSTestDB() error { + for name, db := range pointerDoSFixtures() { + path := filepath.Clean(filepath.Join(w.target, name)) + if err := os.WriteFile(path, db, 0o644); err != nil { + return fmt.Errorf("writing pointer DoS database %s: %w", name, err) + } + } + return nil +} diff --git a/pkg/writer/pointerdos_test.go b/pkg/writer/pointerdos_test.go new file mode 100644 index 00000000..519b4375 --- /dev/null +++ b/pkg/writer/pointerdos_test.go @@ -0,0 +1,598 @@ +package writer + +import ( + "bytes" + "math" + "net/netip" + "os" + "path/filepath" + "strings" + "testing" + + "github.com/oschwald/maxminddb-golang/v2" +) + +// countLeafDecodes walks the fan-out data section the way a reader without +// pointer memoization does, following every pointer and counting how many times +// the shared leaf is decoded. The leaf count is 2**depth and the total number +// of decode operations is Θ(2**depth), which makes the structure a denial of +// service. +func countLeafDecodes(data []byte, offset int) int { + if data[offset]>>5 == 5 { // uint16 leaf + return 1 + } + // Array of size two: a two-byte header followed by two one-byte-payload + // pointers. A pointer's value is ((byte0 & 0x7) << 8) | byte1. + left := int(data[offset+2]&0x7)<<8 | int(data[offset+3]) + right := int(data[offset+4]&0x7)<<8 | int(data[offset+5]) + return countLeafDecodes(data, left) + countLeafDecodes(data, right) +} + +func TestPointerFanOutDataIsExponential(t *testing.T) { + for depth := 1; depth <= 10; depth++ { + data, top := buildPointerFanOutData(depth) + got := countLeafDecodes(data, top) + want := 1 << depth + if got != want { + t.Errorf("depth %d: leaf decoded %d times, want %d", depth, got, want) + } + } +} + +func TestPointerFanOutFixturesMatchCommitted(t *testing.T) { + for name, got := range pointerDoSFixtures() { + //nolint:gosec // name comes from the fixed cases map, not user input. + want, err := os.ReadFile(filepath.Join("..", "..", "test-data", name)) + if err != nil { + t.Fatalf("reading %s: %v", name, err) + } + if !bytes.Equal(got, want) { + t.Errorf("%s differs from the committed file; regenerate with write-test-data", name) + } + } +} + +func TestDecoderLimitBoundaryData(t *testing.T) { + valueTests := []struct { + pointerCount int + wantCount int + wantValues int + }{ + {recommendedValueLimit - 1, 65_535, 65_536}, + {recommendedValueLimit, 65_536, 65_537}, + } + for _, test := range valueTests { + data, top := buildValueLimitData(test.pointerCount) + if top != 1 { + t.Errorf("value boundary top offset = %d, want 1", top) + } + if data[0] != 0xA0 { + t.Errorf("value boundary scalar = %#x, want 0xa0", data[0]) + } + if len(data) != 1+4+test.pointerCount*2 { + t.Errorf( + "value boundary length = %d, want %d", + len(data), + 1+4+test.pointerCount*2, + ) + } + count, headerSize := arrayElementCount(t, data, top) + if count != test.wantCount { + t.Errorf("array element count = %d, want %d", count, test.wantCount) + } + if values := count + 1; values != test.wantValues { + t.Errorf("decoded value count = %d, want %d", values, test.wantValues) + } + for offset := top + headerSize; offset < len(data); offset += 2 { + if target := smallDataPointer(t, data, offset); target != 0 { + t.Fatalf("value boundary pointer at %d targets %d, want 0", offset, target) + } + } + } + + payloadTests := []struct { + smallSize int + wantPayload int + }{ + {32, 2_097_152}, + {33, 2_097_153}, + } + for _, test := range payloadTests { + data, top := buildPayloadLimitData(test.smallSize) + if top <= payloadScalarSize { + t.Errorf("payload boundary top offset = %d, want after large scalar", top) + } + if len(data) <= top { + t.Fatalf("payload boundary data ends before outer array at %d", top) + } + gotPayload := sumPointedPayload(t, data, top) + if gotPayload != test.wantPayload { + t.Errorf("payload boundary total = %d, want %d", gotPayload, test.wantPayload) + } + } +} + +func TestDecoderLimitBoundaryFixturesOpen(t *testing.T) { + tests := []struct { + name string + wantOffset uintptr + }{ + {valueLimitFixtureFilename, 1}, + {valueLimitOverFixtureFilename, 1}, + {payloadLimitFixtureFilename, 65_572}, + {payloadLimitOverFixtureFilename, 65_573}, + } + for _, test := range tests { + t.Run(test.name, func(t *testing.T) { + path := filepath.Clean(filepath.Join("..", "..", "test-data", test.name)) + db, err := maxminddb.Open(path) + if err != nil { + t.Fatalf("opening fixture: %v", err) + } + t.Cleanup(func() { + if err := db.Close(); err != nil { + t.Errorf("closing fixture: %v", err) + } + }) + + result := db.Lookup(netip.MustParseAddr("1.1.1.1")) + if err := result.Err(); err != nil { + t.Fatalf("lookup: %v", err) + } + if !result.Found() { + t.Fatal("lookup did not find a record") + } + if got := result.Offset(); got != test.wantOffset { + t.Errorf("record offset = %d, want %d", got, test.wantOffset) + } + }) + } +} + +func TestDecodePathSharedBudgetData(t *testing.T) { + data, top := buildDecodePathSharedBudgetData() + if top != 3+decodePathSharedKeySize { + t.Fatalf("map offset = %d, want %d", top, 3+decodePathSharedKeySize) + } + if got := len(data); got >= 16*1024 { + t.Errorf("fixture data size = %d, want less than 16 KiB", got) + } + + pos := top + if got, want := data[pos:pos+3], []byte{0xFE, 0x00, 0xE3}; !bytes.Equal(got, want) { + t.Fatalf("map header = %#v, want %#v", got, want) + } + pos += 3 + for i := range decodePathDecoyCount { + if target := smallDataPointer(t, data, pos); target != 0 { + t.Fatalf("decoy key %d points to %d, want 0", i, target) + } + pos += 2 + if got, want := data[pos:pos+2], []byte{0x00, 0x07}; !bytes.Equal(got, want) { + t.Fatalf("decoy value %d = %#v, want %#v", i, got, want) + } + pos += 2 + } + + wantTargetKey := []byte{0x46, 't', 'a', 'r', 'g', 'e', 't'} + if got := data[pos : pos+7]; !bytes.Equal(got, wantTargetKey) { + t.Fatalf("target key = %#v, want %#v", got, wantTargetKey) + } + pos += 7 + if got, want := data[pos:pos+3], []byte{0x5E, 0x0E, 0xDE}; !bytes.Equal(got, want) { + t.Fatalf("selected value header = %#v, want %#v", got, want) + } + + const targetKeySize = len("target") + navigated := decodePathDecoyCount*decodePathSharedKeySize + targetKeySize + if got := navigated + decodePathValueSize; got != fixturePayloadLimit+1 { + t.Errorf("navigated and selected payload = %d, want %d", got, fixturePayloadLimit+1) + } +} + +func TestDecodePathSharedBudgetFixtureIsSemanticallyValid(t *testing.T) { + path := filepath.Clean(filepath.Join("..", "..", "test-data", decodePathBudgetFixtureFilename)) + db, err := maxminddb.Open(path) + if err != nil { + t.Fatalf("opening decode path budget fixture: %v", err) + } + t.Cleanup(func() { + if err := db.Close(); err != nil { + t.Errorf("closing decode path budget fixture: %v", err) + } + }) + + result := db.Lookup(netip.MustParseAddr("1.1.1.1")) + if err := result.Err(); err != nil { + t.Fatalf("lookup: %v", err) + } + // The fixture sits one byte above its 2 MiB example payload budget. The spec + // does not prescribe this threshold, so both outcomes are compliant: the + // reader returns the whole value, or it refuses on its own limit. Assert that + // disjunction. A partial value, a wrong length, or a panic still fails. The + // bytes are checked by TestDecodePathSharedBudgetData and + // TestPointerFanOutFixturesMatchCommitted, which do not depend on the reader. + var value string + err = result.DecodePath(&value, "target") + if err != nil { + if !isReaderResourceLimitError(err) { + t.Fatalf("decoding path: %v", err) + } + t.Logf("reader refused the path decode, which its limits allow: %v", err) + } else if len(value) != decodePathValueSize { + t.Errorf("selected value length = %d, want %d", len(value), decodePathValueSize) + } +} + +func TestMetadataLimitFixtureIsSemanticallyValid(t *testing.T) { + path := filepath.Clean(filepath.Join("..", "..", "test-data", metadataLimitFixtureFilename)) + // The metadata materializes 2,228,190 bytes while opening, above the 2 MiB + // example payload budget used by these fixtures. A reader that bounds metadata + // rejects the file here, which is the behavior this fixture exists to provoke, + // so treat that as a pass and check the sizes only when the reader accepts it. + db, err := maxminddb.Open(path) + if err != nil { + if !isReaderResourceLimitError(err) { + t.Fatalf("opening metadata limit fixture: %v", err) + } + t.Logf("reader refused to open the fixture, which its limits allow: %v", err) + return + } + t.Cleanup(func() { + if err := db.Close(); err != nil { + t.Errorf("closing metadata limit fixture: %v", err) + } + }) + + if got := len(db.Metadata.DatabaseType); got != payloadScalarSize { + t.Errorf("database type length = %d, want %d", got, payloadScalarSize) + } + if got := len(db.Metadata.Languages); got != 33 { + t.Fatalf("language count = %d, want 33", got) + } + for i, language := range db.Metadata.Languages { + if got := len(language); got != payloadScalarSize { + t.Errorf("language %d length = %d, want %d", i, got, payloadScalarSize) + } + } +} + +func isReaderResourceLimitError(err error) bool { + return strings.Contains(err.Error(), "maximum decoded record size") +} + +func TestPointerFanOutFixturesAreSemanticallyValid(t *testing.T) { + tests := []struct { + name string + ipVersion uint + nodeCount uint + depth int + addresses []string + }{ + { + name: pointerDoSFixtureFilename, + ipVersion: 4, + nodeCount: 1, + depth: pointerDoSDepth, + addresses: []string{"0.0.0.0", "203.0.113.9", "255.255.255.255"}, + }, + { + name: pointerDoSIPv6FixtureFilename, + ipVersion: 6, + nodeCount: 97, + depth: pointerDoSDepth, + addresses: []string{ + "0.0.0.0", + "203.0.113.9", + "255.255.255.255", + "::", + "2001:db8::1", + "ffff:ffff:ffff:ffff:ffff:ffff:ffff:ffff", + }, + }, + { + name: pointerValueLimitFixtureFilename, + ipVersion: 4, + nodeCount: 1, + depth: 15, + addresses: []string{"1.1.1.1"}, + }, + } + + for _, test := range tests { + t.Run(test.name, func(t *testing.T) { + path := filepath.Clean(filepath.Join("..", "..", "test-data", test.name)) + db, err := maxminddb.Open(path) + if err != nil { + t.Fatalf("opening fixture: %v", err) + } + t.Cleanup(func() { + if err := db.Close(); err != nil { + t.Errorf("closing fixture: %v", err) + } + }) + + if got := db.Metadata.BinaryFormatMajorVersion; got != 2 { + t.Errorf("binary format major version = %d, want 2", got) + } + if got := db.Metadata.BinaryFormatMinorVersion; got != 0 { + t.Errorf("binary format minor version = %d, want 0", got) + } + if got := db.Metadata.BuildEpoch; got != pointerDoSBuildEpoch { + t.Errorf("build epoch = %d, want %d", got, pointerDoSBuildEpoch) + } + if got := db.Metadata.DatabaseType; got != "Test" { + t.Errorf("database type = %q, want %q", got, "Test") + } + if got := db.Metadata.IPVersion; got != test.ipVersion { + t.Errorf("IP version = %d, want %d", got, test.ipVersion) + } + if got := db.Metadata.NodeCount; got != test.nodeCount { + t.Errorf("node count = %d, want %d", got, test.nodeCount) + } + if got := db.Metadata.RecordSize; got != 24 { + t.Errorf("record size = %d, want 24", got) + } + + // Lookups stop at the outer record. Do not call Decode here: expanding + // the fan-out is deliberately hostile in an unprotected reader. + var outerOffset uintptr + for i, address := range test.addresses { + result := db.Lookup(netip.MustParseAddr(address)) + if err := result.Err(); err != nil { + t.Fatalf("looking up %s: %v", address, err) + } + if !result.Found() { + t.Fatalf("lookup for %s did not find a record", address) + } + if i == 0 { + outerOffset = result.Offset() + } else if got := result.Offset(); got != outerOffset { + t.Errorf("lookup for %s returned offset %d, want %d", address, got, outerOffset) + } + } + + // The reader independently validates the metadata and search tree. Walk + // only one branch per level to validate the data topology in bounded time. + raw, err := os.ReadFile(path) + if err != nil { + t.Fatalf("reading fixture: %v", err) + } + dataStart := dataSectionStart(t, db, raw) + if uint64(outerOffset) > uint64(math.MaxInt) { + t.Fatalf("outer record offset %d overflows int", outerOffset) + } + outerOffsetInt := int(outerOffset) + if outerOffsetInt >= len(raw)-dataStart { + t.Fatalf("outer record offset %d is outside the data section", outerOffset) + } + verifyPointerFanOutRecord(t, raw[dataStart:], outerOffsetInt, test.depth) + }) + } +} + +// verifyPointerFanOutRecord follows one of the two equal pointers at each +// level. Its work is linear in depth and never expands the hostile structure. +func verifyPointerFanOutRecord(t *testing.T, data []byte, offset, depth int) { + t.Helper() + for level := depth; level > 0; level-- { + if offset < 0 || offset+6 > len(data) { + t.Fatalf("level %d array at offset %d is outside the data section", level, offset) + } + if data[offset] != 0x02 || data[offset+1] != 0x04 { + t.Fatalf("level %d at offset %d is not a two-element array", level, offset) + } + left := smallDataPointer(t, data, offset+2) + right := smallDataPointer(t, data, offset+4) + if left != right { + t.Fatalf("level %d pointers differ: left %d, right %d", level, left, right) + } + if left >= offset { + t.Fatalf( + "level %d pointer target %d does not precede array offset %d", + level, + left, + offset, + ) + } + offset = left + } + if offset != 0 { + t.Fatalf("leaf offset = %d, want 0", offset) + } + if data[offset] != 0xA0 { + t.Fatalf("leaf control byte = %#x, want 0xa0", data[offset]) + } +} + +// scalarPayloadSize returns the declared payload length of the string or bytes +// scalar at offset, reading only its control bytes. +func scalarPayloadSize(t *testing.T, data []byte, offset int) int { + t.Helper() + if offset < 0 || offset+3 > len(data) { + t.Fatalf("scalar at offset %d is outside the data section", offset) + } + control := data[offset] + major := control >> 5 + if major != scalarTypeString && major != scalarTypeBytes { + t.Fatalf("value at offset %d has major type %d, want a string or bytes", offset, major) + } + size := int(control & 0x1F) + switch { + case size <= 28: + return size + case size == 29: + return int(data[offset+1]) + 29 + case size == 30: + return (int(data[offset+1])<<8 | int(data[offset+2])) + 285 + default: + t.Fatalf("scalar at offset %d uses size code %d, which this test cannot read", offset, size) + return 0 + } +} + +// sumPointedPayload adds the payload referenced by each element of the array at +// top. This is the per-occurrence total used by the flat payload-budget fixtures; +// it does not imply that a memoizing reader materializes each target again. The +// function reads the generated bytes rather than recomputing the builder's +// arithmetic, so a change to the encoding or element count fails the assertion. +func sumPointedPayload(t *testing.T, data []byte, top int) int { + t.Helper() + count, headerSize := arrayElementCount(t, data, top) + total := 0 + elem := top + headerSize + for range count { + total += scalarPayloadSize(t, data, smallDataPointer(t, data, elem)) + elem += 2 + } + return total +} + +func arrayElementCount(t *testing.T, data []byte, offset int) (count, headerSize int) { + t.Helper() + if offset < 0 || offset+2 > len(data) || data[offset]>>5 != 0 || data[offset+1] != 4 { + t.Fatalf("value at offset %d is not an array", offset) + } + sizeCode := int(data[offset] & 0x1F) + switch { + case sizeCode <= 28: + return sizeCode, 2 + case sizeCode == 29: + if offset+3 > len(data) { + t.Fatalf("size-29 array header at offset %d is truncated", offset) + } + return int(data[offset+2]) + 29, 3 + case sizeCode == 30: + if offset+4 > len(data) { + t.Fatalf("size-30 array header at offset %d is truncated", offset) + } + return (int(data[offset+2])<<8 | int(data[offset+3])) + 285, 4 + default: + if offset+5 > len(data) { + t.Fatalf("size-31 array header at offset %d is truncated", offset) + } + return (int(data[offset+2])<<16 | int(data[offset+3])<<8 | int(data[offset+4])) + + minimumSizeCode31, 5 + } +} + +func smallDataPointer(t *testing.T, data []byte, offset int) int { + t.Helper() + if offset < 0 || offset+2 > len(data) { + t.Fatalf("pointer at offset %d is outside the data section", offset) + } + control := data[offset] + if control>>5 != 1 || control&0x18 != 0 { + t.Fatalf("value at offset %d is not a two-byte data pointer", offset) + } + return int(control&0x7)<<8 | int(data[offset+1]) +} + +// dataSectionStart returns the offset in raw where the data section begins, +// after the search tree and its separator, validating the tree dimensions fit +// an int. +func dataSectionStart(t *testing.T, db *maxminddb.Reader, raw []byte) int { + t.Helper() + recordSizeQuarter := uint64(db.Metadata.RecordSize / 4) + if recordSizeQuarter == 0 || + uint64(db.Metadata.NodeCount) > uint64(math.MaxInt-dataSeparatorSize)/recordSizeQuarter { + t.Fatalf("search tree dimensions overflow int: %d nodes at %d bits per record", + db.Metadata.NodeCount, db.Metadata.RecordSize) + } + searchTreeSize64 := uint64(db.Metadata.NodeCount) * recordSizeQuarter + //nolint:gosec // G115: searchTreeSize64 is checked against math.MaxInt above. + dataStart := int(searchTreeSize64) + dataSeparatorSize + if dataStart > len(raw) { + t.Fatalf("data section starts at %d, beyond file size %d", dataStart, len(raw)) + } + return dataStart +} + +// TestPayloadAmplificationFixturesAreSemanticallyValid independently validates +// each committed payload fixture: a shared 65,535-byte scalar at data-section +// offset 0 and an outer array whose elements all point to it. It inspects the +// raw data section without expanding the array. +func TestPayloadAmplificationFixturesAreSemanticallyValid(t *testing.T) { + tests := []struct { + name string + scalarType byte // MMDB major type: 2 is string, 4 is bytes + pointerCount int + }{ + {payloadDoSFixtureFilename, scalarTypeBytes, payloadPointerCount}, + {worstCasePayloadFixtureFilename, scalarTypeBytes, worstCasePointerCount}, + {stringPayloadFixtureFilename, scalarTypeString, payloadPointerCount}, + } + + for _, test := range tests { + t.Run(test.name, func(t *testing.T) { + path := filepath.Clean(filepath.Join("..", "..", "test-data", test.name)) + db, err := maxminddb.Open(path) + if err != nil { + t.Fatalf("opening fixture: %v", err) + } + t.Cleanup(func() { + if err := db.Close(); err != nil { + t.Errorf("closing fixture: %v", err) + } + }) + + if got := db.Metadata.IPVersion; got != 4 { + t.Errorf("IP version = %d, want 4", got) + } + + // The lookup finds the outer array. Do not Decode: expanding it is + // deliberately hostile in an unprotected reader. + result := db.Lookup(netip.MustParseAddr("1.1.1.1")) + if err := result.Err(); err != nil { + t.Fatalf("lookup: %v", err) + } + if !result.Found() { + t.Fatal("lookup did not find a record") + } + offset := result.Offset() + if uint64(offset) > uint64(math.MaxInt) { + t.Fatalf("record offset %d overflows int", offset) + } + //nolint:gosec // G115: offset is checked against math.MaxInt above. + outerOffset := int(offset) + + raw, err := os.ReadFile(path) + if err != nil { + t.Fatalf("reading fixture: %v", err) + } + data := raw[dataSectionStart(t, db, raw):] + + // The shared scalar sits at offset 0. Size code 30 means its length is + // the next two bytes plus 285. + wantControl := (test.scalarType << 5) | 30 + if got := data[0]; got != wantControl { + t.Errorf("scalar control byte = %#x, want %#x", got, wantControl) + } + if got := (int(data[1])<<8 | int(data[2])) + 285; got != payloadScalarSize { + t.Errorf("scalar length = %d, want %d", got, payloadScalarSize) + } + + // The outer record is an extended array (0x1E, 0x04) with a size-30 + // two-byte element count. + if outerOffset+4 > len(data) { + t.Fatalf("array header at %d is outside the data section", outerOffset) + } + if data[outerOffset] != 0x1E || data[outerOffset+1] != 0x04 { + t.Fatalf("outer record at %d is not an extended array", outerOffset) + } + count := (int(data[outerOffset+2])<<8 | int(data[outerOffset+3])) + 285 + if count != test.pointerCount { + t.Errorf("array element count = %d, want %d", count, test.pointerCount) + } + + // Every element is a two-byte pointer to the shared value at offset 0. + elem := outerOffset + 4 + for i := range test.pointerCount { + if target := smallDataPointer(t, data, elem); target != 0 { + t.Fatalf("array element %d points to offset %d, want 0", i, target) + } + elem += 2 + } + }) + } +} diff --git a/pkg/writer/rawmmdb.go b/pkg/writer/rawmmdb.go index b91b34c2..2fd01326 100644 --- a/pkg/writer/rawmmdb.go +++ b/pkg/writer/rawmmdb.go @@ -6,11 +6,18 @@ package writer // large-size (case-31) encoding and complete database builders for crafting // intentionally malformed MMDB files that cannot be created through mmdbwriter. -import "encoding/binary" +import ( + "encoding/binary" + "fmt" +) const ( - metadataMarker = "\xab\xcd\xefMaxMind.com" - dataSeparatorSize = 16 + metadataMarker = "\xab\xcd\xefMaxMind.com" + dataSeparatorSize = 16 + maximumSizeCode30 = 65820 + minimumSizeCode31 = maximumSizeCode30 + 1 + maximumDataStructureSize = minimumSizeCode31 + (1 << 24) - 1 + maximum24BitSearchTreeValue = 1<<24 - 1 ) var ( @@ -49,17 +56,42 @@ var ( } ) -// writeMap writes a map control byte (type 7) for sizes <= 28. +// writeMap writes a map control byte (type 7). Sizes above 28 use the extended +// size forms, which take one to three more bytes. func writeMap(buf []byte, size int) int { - buf[0] = (7 << 5) | byte(size&0x1f) - return 1 + if size < 0 || size > maximumDataStructureSize { + panic(fmt.Sprintf( + "map size %d is outside the supported range 0..%d", + size, + maximumDataStructureSize, + )) + } + switch { + case size <= 28: + buf[0] = (7 << 5) | byte(size&0x1F) + return 1 + case size <= 284: + buf[0] = (7 << 5) | 29 + buf[1] = byte(size - 29) + return 2 + case size <= maximumSizeCode30: + buf[0] = (7 << 5) | 30 + encoded := size - 285 + buf[1] = byte((encoded >> 8) & 0xFF) + buf[2] = byte(encoded & 0xFF) + return 3 + default: + return writeLargeMap(buf, uint32(size)) + } } -// writeString writes a string value (type 2). +// writeString writes a string value (type 2). It delegates the control byte to +// writeScalar so strings longer than 28 bytes get the extended size forms +// instead of a truncated size. func writeString(buf []byte, s string) int { - buf[0] = (2 << 5) | byte(len(s)&0x1f) - copy(buf[1:], s) - return 1 + len(s) + pos := writeScalar(buf, len(s), scalarTypeString) + copy(buf[pos-len(s):pos], s) + return pos } // writeUint16 writes a uint16 value (type 5, 2 bytes). @@ -92,7 +124,15 @@ func writeMetaKey(buf []byte, key string) int { // writeLargeArray writes an array control byte (extended type 11) with // case-31 size encoding for sizes > 65820. func writeLargeArray(buf []byte, size uint32) int { - adjusted := size - 65821 + if size < minimumSizeCode31 || size > maximumDataStructureSize { + panic(fmt.Sprintf( + "array size %d is outside the case-31 range %d..%d", + size, + minimumSizeCode31, + maximumDataStructureSize, + )) + } + adjusted := size - minimumSizeCode31 buf[0] = (0 << 5) | 31 // extended type, size = case 31 buf[1] = 4 // extended type: 7 + 4 = 11 (array) buf[2] = byte((adjusted >> 16) & 0xFF) @@ -104,7 +144,15 @@ func writeLargeArray(buf []byte, size uint32) int { // writeLargeMap writes a map control byte (type 7) with case-31 size // encoding for sizes > 65820. func writeLargeMap(buf []byte, size uint32) int { - adjusted := size - 65821 + if size < minimumSizeCode31 || size > maximumDataStructureSize { + panic(fmt.Sprintf( + "map size %d is outside the case-31 range %d..%d", + size, + minimumSizeCode31, + maximumDataStructureSize, + )) + } + adjusted := size - minimumSizeCode31 buf[0] = (7 << 5) | 31 // type 7 (map), size = case 31 buf[1] = byte((adjusted >> 16) & 0xFF) buf[2] = byte((adjusted >> 8) & 0xFF) @@ -125,9 +173,16 @@ func writeSearchTree(buf []byte, recordValue uint32) int { return writeSearchTreeRecords(buf, recordValue, recordValue) } -// writeSearchTreeRecords writes a 1-node search tree with 24-bit records -// where the left and right records can hold different values. +// writeSearchTreeRecords writes one search-tree node with 24-bit records where +// the left and right records can hold different values. func writeSearchTreeRecords(buf []byte, leftRecord, rightRecord uint32) int { + if leftRecord > maximum24BitSearchTreeValue || rightRecord > maximum24BitSearchTreeValue { + panic(fmt.Sprintf( + "search-tree record values %d and %d must fit in 24 bits", + leftRecord, + rightRecord, + )) + } buf[0] = byte((leftRecord >> 16) & 0xFF) buf[1] = byte((leftRecord >> 8) & 0xFF) buf[2] = byte(leftRecord & 0xFF) @@ -141,6 +196,7 @@ func writeMetadataBlockWithKeyOrder( buf []byte, nodeCount uint32, buildEpoch uint64, + ipVersion uint16, keys []string, ) int { pos := 0 @@ -156,7 +212,7 @@ func writeMetadataBlockWithKeyOrder( "build_epoch": func(b []byte) int { return writeUint64(b, buildEpoch) }, "database_type": func(b []byte) int { return writeString(b, "Test") }, "description": func(b []byte) int { return writeMap(b, 0) }, - "ip_version": func(b []byte) int { return writeUint16(b, 4) }, + "ip_version": func(b []byte) int { return writeUint16(b, ipVersion) }, "languages": writeEmptyArray, "node_count": func(b []byte) int { return writeUint32(b, nodeCount) }, "record_size": func(b []byte) int { return writeUint16(b, 24) }, @@ -177,7 +233,7 @@ func writeMetadataBlockWithKeyOrder( // writeMetadataBlock writes the metadata marker followed by a standard // metadata map with the given parameters. func writeMetadataBlock(buf []byte, nodeCount uint32, buildEpoch uint64) int { - return writeMetadataBlockWithKeyOrder(buf, nodeCount, buildEpoch, metadataKeysStandard) + return writeMetadataBlockWithKeyOrder(buf, nodeCount, buildEpoch, 4, metadataKeysStandard) } func buildSimpleDB(metadataWriter func([]byte, uint32, uint64) int) []byte { @@ -268,6 +324,7 @@ func writeMetadataBlockEmptyMapLast(buf []byte, nodeCount uint32, buildEpoch uin buf, nodeCount, buildEpoch, + 4, metadataKeysEmptyMapLast, ) } @@ -279,6 +336,7 @@ func writeMetadataBlockEmptyArrayLast(buf []byte, nodeCount uint32, buildEpoch u buf, nodeCount, buildEpoch, + 4, metadataKeysEmptyArrayLast, ) } diff --git a/pkg/writer/rawmmdb_test.go b/pkg/writer/rawmmdb_test.go new file mode 100644 index 00000000..0386a5e6 --- /dev/null +++ b/pkg/writer/rawmmdb_test.go @@ -0,0 +1,190 @@ +package writer + +import ( + "bytes" + "fmt" + "strings" + "testing" +) + +// TestWriteMapControlByte covers the direct form and all three extended size +// forms. Current fixtures exercise the one- and three-byte map headers, so +// explicit cases preserve the two- and four-byte boundaries. +func TestWriteMapControlByte(t *testing.T) { + tests := []struct { + size int + want []byte + }{ + {0, []byte{0xE0}}, + {28, []byte{0xFC}}, + {29, []byte{0xFD, 0x00}}, + {284, []byte{0xFD, 0xFF}}, + {285, []byte{0xFE, 0x00, 0x00}}, + {512, []byte{0xFE, 0x00, 0xE3}}, + {65820, []byte{0xFE, 0xFF, 0xFF}}, + {65821, []byte{0xFF, 0x00, 0x00, 0x00}}, + {1_000_000, []byte{0xFF, 0x0E, 0x41, 0x23}}, + {maximumDataStructureSize, []byte{0xFF, 0xFF, 0xFF, 0xFF}}, + } + for _, test := range tests { + buf := make([]byte, 8) + n := writeMap(buf, test.size) + if got := buf[:n]; !bytes.Equal(got, test.want) { + t.Errorf("writeMap(%d) = %#v, want %#v", test.size, got, test.want) + } + } +} + +// TestWriteStringControlByte covers the boundary where the size no longer fits +// the control byte. Every current caller passes a short metadata key, so nothing +// else exercises the extended forms. +func TestWriteStringControlByte(t *testing.T) { + tests := []struct { + size int + wantHeader []byte + }{ + {0, []byte{0x40}}, + {2, []byte{0x42}}, + {28, []byte{0x5C}}, + {29, []byte{0x5D, 0x00}}, + {284, []byte{0x5D, 0xFF}}, + {285, []byte{0x5E, 0x00, 0x00}}, + {65820, []byte{0x5E, 0xFF, 0xFF}}, + {65821, []byte{0x5F, 0x00, 0x00, 0x00}}, + {1_000_000, []byte{0x5F, 0x0E, 0x41, 0x23}}, + } + for _, test := range tests { + value := strings.Repeat("a", test.size) + buf := make([]byte, test.size+8) + n := writeString(buf, value) + want := append(append([]byte{}, test.wantHeader...), value...) + if got := buf[:n]; !bytes.Equal(got, want) { + t.Errorf("writeString(%d bytes) = %#v, want %#v", + test.size, got[:min(n, 4)], want[:min(len(want), 4)]) + } + } +} + +func TestWriteArrayHeader(t *testing.T) { + tests := []struct { + size int + want []byte + }{ + {0, []byte{0x00, 0x04}}, + {28, []byte{0x1C, 0x04}}, + {29, []byte{0x1D, 0x04, 0x00}}, + {284, []byte{0x1D, 0x04, 0xFF}}, + {285, []byte{0x1E, 0x04, 0x00, 0x00}}, + {65820, []byte{0x1E, 0x04, 0xFF, 0xFF}}, + {65821, []byte{0x1F, 0x04, 0x00, 0x00, 0x00}}, + {1_000_000, []byte{0x1F, 0x04, 0x0E, 0x41, 0x23}}, + {maximumDataStructureSize, []byte{0x1F, 0x04, 0xFF, 0xFF, 0xFF}}, + } + for _, test := range tests { + buf := make([]byte, 8) + n := writeArrayHeader(buf, test.size) + if got := buf[:n]; !bytes.Equal(got, test.want) { + t.Errorf("writeArrayHeader(%d) = %#v, want %#v", test.size, got, test.want) + } + } +} + +func TestWriteSearchTreeRecordsMaximum(t *testing.T) { + buf := make([]byte, 6) + n := writeSearchTreeRecords(buf, maximum24BitSearchTreeValue, maximum24BitSearchTreeValue) + want := []byte{0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF} + if got := buf[:n]; !bytes.Equal(got, want) { + t.Errorf("writeSearchTreeRecords() = %#v, want %#v", got, want) + } +} + +func TestRawWritersRejectInvalidInput(t *testing.T) { + tests := []struct { + name string + call func() + wantPanic string + }{ + { + "negative map size", + func() { writeMap(make([]byte, 8), -1) }, + "map size -1 is outside the supported range", + }, + { + "oversized map", + func() { writeMap(make([]byte, 8), maximumDataStructureSize+1) }, + "map size 16843037 is outside the supported range", + }, + { + "undersized large map", + func() { writeLargeMap(make([]byte, 8), maximumSizeCode30) }, + "map size 65820 is outside the case-31 range", + }, + { + "oversized large map", + func() { writeLargeMap(make([]byte, 8), maximumDataStructureSize+1) }, + "map size 16843037 is outside the case-31 range", + }, + { + "negative array size", + func() { writeArrayHeader(make([]byte, 8), -1) }, + "array size -1 is outside the supported range", + }, + { + "oversized array", + func() { writeArrayHeader(make([]byte, 8), maximumDataStructureSize+1) }, + "array size 16843037 is outside the supported range", + }, + { + "undersized large array", + func() { writeLargeArray(make([]byte, 8), maximumSizeCode30) }, + "array size 65820 is outside the case-31 range", + }, + { + "oversized large array", + func() { writeLargeArray(make([]byte, 8), maximumDataStructureSize+1) }, + "array size 16843037 is outside the case-31 range", + }, + { + "negative scalar size", + func() { writeScalar(make([]byte, 8), -1, scalarTypeBytes) }, + "scalar size -1 is outside the supported range", + }, + { + "oversized scalar", + func() { writeScalar(make([]byte, 8), maximumDataStructureSize+1, scalarTypeBytes) }, + "scalar size 16843037 is outside the supported range", + }, + { + "invalid scalar type", + func() { writeScalar(make([]byte, 8), 0, 8) }, + "unsupported scalar type 8", + }, + { + "oversized left search-tree record", + func() { + writeSearchTreeRecords(make([]byte, 8), maximum24BitSearchTreeValue+1, 0) + }, + "search-tree record values 16777216 and 0 must fit in 24 bits", + }, + { + "oversized right search-tree record", + func() { + writeSearchTreeRecords(make([]byte, 8), 0, maximum24BitSearchTreeValue+1) + }, + "search-tree record values 0 and 16777216 must fit in 24 bits", + }, + } + for _, test := range tests { + t.Run(test.name, func(t *testing.T) { + defer func() { + got := recover() + if got == nil { + t.Error("call did not panic") + } else if message := fmt.Sprint(got); !strings.Contains(message, test.wantPanic) { + t.Errorf("panic = %q, want it to contain %q", message, test.wantPanic) + } + }() + test.call() + }) + } +} diff --git a/test-data/MaxMind-DB-test-decode-path-shared-budget.mmdb b/test-data/MaxMind-DB-test-decode-path-shared-budget.mmdb new file mode 100644 index 00000000..6e85ac07 Binary files /dev/null and b/test-data/MaxMind-DB-test-decode-path-shared-budget.mmdb differ diff --git a/test-data/MaxMind-DB-test-decoder-payload-limit-over.mmdb b/test-data/MaxMind-DB-test-decoder-payload-limit-over.mmdb new file mode 100644 index 00000000..5d8d82b8 Binary files /dev/null and b/test-data/MaxMind-DB-test-decoder-payload-limit-over.mmdb differ diff --git a/test-data/MaxMind-DB-test-decoder-payload-limit.mmdb b/test-data/MaxMind-DB-test-decoder-payload-limit.mmdb new file mode 100644 index 00000000..bb09c158 Binary files /dev/null and b/test-data/MaxMind-DB-test-decoder-payload-limit.mmdb differ diff --git a/test-data/MaxMind-DB-test-decoder-value-limit-over.mmdb b/test-data/MaxMind-DB-test-decoder-value-limit-over.mmdb new file mode 100644 index 00000000..9310644d Binary files /dev/null and b/test-data/MaxMind-DB-test-decoder-value-limit-over.mmdb differ diff --git a/test-data/MaxMind-DB-test-decoder-value-limit-pointer-heavy.mmdb b/test-data/MaxMind-DB-test-decoder-value-limit-pointer-heavy.mmdb new file mode 100644 index 00000000..f3ee81d6 Binary files /dev/null and b/test-data/MaxMind-DB-test-decoder-value-limit-pointer-heavy.mmdb differ diff --git a/test-data/MaxMind-DB-test-decoder-value-limit.mmdb b/test-data/MaxMind-DB-test-decoder-value-limit.mmdb new file mode 100644 index 00000000..84cde977 Binary files /dev/null and b/test-data/MaxMind-DB-test-decoder-value-limit.mmdb differ diff --git a/test-data/MaxMind-DB-test-metadata-payload-limit.mmdb b/test-data/MaxMind-DB-test-metadata-payload-limit.mmdb new file mode 100644 index 00000000..86f694d1 Binary files /dev/null and b/test-data/MaxMind-DB-test-metadata-payload-limit.mmdb differ diff --git a/test-data/MaxMind-DB-test-payload-amplification-dos-string.mmdb b/test-data/MaxMind-DB-test-payload-amplification-dos-string.mmdb new file mode 100644 index 00000000..0dee394d Binary files /dev/null and b/test-data/MaxMind-DB-test-payload-amplification-dos-string.mmdb differ diff --git a/test-data/MaxMind-DB-test-payload-amplification-dos-worst-case.mmdb b/test-data/MaxMind-DB-test-payload-amplification-dos-worst-case.mmdb new file mode 100644 index 00000000..49316a6e Binary files /dev/null and b/test-data/MaxMind-DB-test-payload-amplification-dos-worst-case.mmdb differ diff --git a/test-data/MaxMind-DB-test-payload-amplification-dos.mmdb b/test-data/MaxMind-DB-test-payload-amplification-dos.mmdb new file mode 100644 index 00000000..94728552 Binary files /dev/null and b/test-data/MaxMind-DB-test-payload-amplification-dos.mmdb differ diff --git a/test-data/MaxMind-DB-test-pointer-decoder-dos-ipv6.mmdb b/test-data/MaxMind-DB-test-pointer-decoder-dos-ipv6.mmdb new file mode 100644 index 00000000..9181a06c Binary files /dev/null and b/test-data/MaxMind-DB-test-pointer-decoder-dos-ipv6.mmdb differ diff --git a/test-data/MaxMind-DB-test-pointer-decoder-dos.mmdb b/test-data/MaxMind-DB-test-pointer-decoder-dos.mmdb new file mode 100644 index 00000000..7798ab25 Binary files /dev/null and b/test-data/MaxMind-DB-test-pointer-decoder-dos.mmdb differ diff --git a/test-data/README.md b/test-data/README.md index 4465d5b4..0307ae26 100644 --- a/test-data/README.md +++ b/test-data/README.md @@ -32,6 +32,83 @@ broken, and exploited functionality simply not available in the go mmdbwriter: - GeoIP2-City-Test-Invalid-Node-Count.mmdb - maps-with-pointers.raw +## Denial of service test data + +Some files in this directory are hostile by design. Each one is structurally +well-formed, but a reader's resource policy may still refuse to open or decode +it. A reader without resource controls can use far more CPU or memory than the +file size suggests. Use these files to test the guidance in the +[Reader Resource Limits](../MaxMind-DB-spec.md#reader-resource-limits) section +of the specification. + +These files do not define an exact boundary for the recommended nesting depth. +Use a reader-level test to verify the exact boundary that implementation uses. + +Do not decode these files without time and memory limits on the process. The +worst-case file is about 192 KiB and describes about 4 GiB of repeated data. A +test harness without resource controls that walks this whole directory and +decodes every data entry will hang or run out of memory. + +| File | Behavior under test | +| --------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| MaxMind-DB-test-pointer-decoder-dos.mmdb | A depth-40 pointer fan-out. An unprotected decoder performs 2\*\*40 leaf decodes from 451 bytes. | +| MaxMind-DB-test-pointer-decoder-dos-ipv6.mmdb | The same fan-out in a conventional IPv6 database that maps the whole address space to the data entry. | +| MaxMind-DB-test-payload-amplification-dos.mmdb | 8,192 pointers to one 65,535-byte value of type `bytes`. A reader that copies each target materializes about 512 MiB. | +| MaxMind-DB-test-payload-amplification-dos-string.mmdb | The same shape with a UTF-8 string value, the type most bindings copy into a native string. | +| MaxMind-DB-test-payload-amplification-dos-worst-case.mmdb | 65,535 pointers to that value. The data entry decodes to 65,536 values and meets the recommended value limit. Copying every occurrence materializes about 4 GiB. | + +The next table lists example boundary fixtures. Its results assume the flat +value accounting rule and the 2 MiB payload limit described below. + +| File | Expected result | +| ------------------------------------------------------ | ------------------------------------------------------------------------------- | +| MaxMind-DB-test-decoder-value-limit.mmdb | Accept. 65,536 decoded values, exactly the limit. | +| MaxMind-DB-test-decoder-value-limit-over.mmdb | Reject. 65,537 decoded values. | +| MaxMind-DB-test-decoder-value-limit-pointer-heavy.mmdb | Accept. 65,535 values reached through a depth-15 fan-out. | +| MaxMind-DB-test-decoder-payload-limit.mmdb | Accept. 2,097,152 payload bytes, exactly 2 MiB. | +| MaxMind-DB-test-decoder-payload-limit-over.mmdb | Reject. 2,097,153 payload bytes. | +| MaxMind-DB-test-metadata-payload-limit.mmdb | Reject at open. The metadata alone materializes 2,228,190 bytes. | +| MaxMind-DB-test-decode-path-shared-budget.mmdb | Reject a path lookup. Navigation plus the selected value costs 2,097,153 bytes. | + +### The value-count boundary is a recommendation + +The spec recommends a limit of 65,536 decoded values. It does not require a +single payload limit, because the right method and limit depend on the reader's +language and API. + +The payload files use 2 MiB (2\*\*21 bytes), matching the current companion pull +request heads for +[libmaxminddb](https://github.com/maxmind/libmaxminddb/pull/479), +[Go](https://github.com/oschwald/maxminddb-golang/pull/233), +[Python](https://github.com/maxmind/MaxMind-DB-Reader-python/pull/439), +[Ruby](https://github.com/maxmind/MaxMind-DB-Reader-ruby/pull/235), +[PHP](https://github.com/maxmind/MaxMind-DB-Reader-php/pull/281), +[Java](https://github.com/maxmind/MaxMind-DB-Reader-java/pull/442), and +[.NET](https://github.com/maxmind/MaxMind-DB-Reader-dotnet/pull/355). In +libmaxminddb, 2 MiB is the default and can be changed at build time; the current +Go, Python, Ruby, PHP, Java, and .NET pull request heads use a fixed 2 MiB +limit. This shared test boundary is not a file format requirement. A reader that +picks a different payload limit or uses an equivalent strategy is still +compliant. Treat these files as examples of the attack shape and adjust the +expected boundary to the reader's policy. + +### The payload files assume per-occurrence accounting + +The payload boundary files point 32 of their 33 elements at one shared value. +They separate accept from reject only for a reader that charges every decoded +occurrence. + +A reader that safely memoizes pointer targets can materialize the shared value +one time, about 64 KiB, and accept both files. That behavior is compliant. The +result can still contain many logical references to the shared value. A binding, +serializer, or conversion to owning values that copies each occurrence should +apply its own bound. The amplification files above can test that downstream +boundary. + +The pointer fan-out files test a reader without safe reuse or another work +bound. The worst-case payload file tests a reader whose only defense is the +decoded-value count. + ## Usage ```