Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
a6c79bd
Document reader resource limits for the pointer fan-out DoS
oschwald Aug 25, 2026
f9493be
Add pointer fan-out DoS test databases
oschwald Aug 25, 2026
5986edb
Add payload amplification DoS test databases
oschwald Aug 26, 2026
3056e15
Add decoder resource limit boundary fixtures
oschwald Aug 26, 2026
b20fa64
Add cumulative decode path budget fixture
oschwald Aug 29, 2026
6466330
Share one fixture map between the writer and the test
oschwald Aug 31, 2026
e7ab6b9
Let the fixture tests accept a reader that enforces limits
oschwald Aug 31, 2026
254545a
Document the denial of service test data
oschwald Aug 31, 2026
0cf3b51
Encode sizes above 28 correctly in the raw writers
oschwald Aug 31, 2026
6c476b1
Ignore the root-level write-test-data binary
oschwald Aug 31, 2026
aa8a9cc
Harden raw MMDB fixture encoders
oschwald Sep 1, 2026
7de1b2a
Strengthen resource limit fixture validation
oschwald Sep 1, 2026
c24bef7
Define decoded value accounting guidance
oschwald Sep 1, 2026
8262fb5
Clarify decode and materialization boundaries
oschwald Sep 1, 2026
bc1fa6c
Remove unpublished advisory reference
oschwald Sep 1, 2026
4ef895a
Warn about hostile MMDB test fixtures
oschwald Sep 1, 2026
150f2af
Document resource limit fixture policies
oschwald Sep 1, 2026
a93cc69
Polish resource limit specification wording
oschwald Sep 1, 2026
42c284e
Improve raw encoder boundary coverage
oschwald Sep 1, 2026
694acee
Clarify resource limit fixture behavior
oschwald Sep 1, 2026
0f06955
Clarify example payload limit terminology
oschwald Sep 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -2,4 +2,5 @@ docs/.hugo_build.lock
docs/public/
*.swp
/cmd/write-test-data/write-test-data
/write-test-data
.lycheecache
135 changes: 135 additions & 0 deletions MaxMind-DB-spec.md
Original file line number Diff line number Diff line change
Expand Up @@ -310,6 +310,10 @@ A pointer to another part of the data section's address space. The pointer will
point to the beginning of a field. It is illegal for a pointer to point to
another pointer.

Because several pointers can share one target, following pointers naively can
decode far more data than a record's size suggests. See
[Reader Resource Limits](#reader-resource-limits).

Pointer values start from the beginning of the data section, _not_ the beginning
of the file. Pointers in the metadata start from the beginning of the metadata
section.
Expand Down Expand Up @@ -537,6 +541,137 @@ are ignored.
This means that we are limited to 4GB of address space for pointers, so the data
section size for the database is limited to 4GB.

## Reader Resource Limits

A crafted or corrupt data section can make a reader use far more time and memory
than a data entry's encoded size suggests. A value can be a pointer, and several
pointers can share one target. A small data section can therefore describe a
structure that is huge, or effectively infinite, when fully expanded.

A reader should apply resource controls to operations that decode
attacker-controlled values. The controls should bound the reader's time and
memory before it allocates, traverses, or copies an attacker-chosen amount of
data. The examples in this section are implementation guidance. They do not
define whether an encoding conforms to the file format.

A caller-requested decode operation can include a lookup, path selection and
decoding of the selected value, or metadata decoding while opening a database. A
reader that validates a whole database can apply a bound to each data entry if
the work outside those entries is also bounded.

A reader can apply the depth, value-count, and payload limits below. It can
instead combine them or use another strategy that bounds the same risks, such
as:

- Safe memoization of pointer targets, including cycle handling.
- Schema-directed decoding into a destination that is finite and nonrecursive,
and that skips unknown values without following pointers.
- A weighted work budget.

A reader that delegates decoding to caller-defined code should document where
that responsibility transfers.

A finite, nonrecursive destination can provide an inherent bound when its
recognized fields have bounded shapes and unknown values are skipped without
following their pointers. A specific typed destination does not provide that
bound by itself. It can still contain recursive types, attacker-sized
collections, repeated recognized fields, or dynamically shaped values. A
schema-aware reader can apply tighter semantic limits where it knows a
collection's valid size is small.

### Maximum Data Structure Depth

A reader that uses a nesting counter should limit how deeply it decodes one data
entry. The depth increases by one each time the reader enters a map or an array,
or follows a pointer. If the depth exceeds 512, the reader should stop and
reject the data entry.

This limit stops unbounded recursion, including a pointer cycle and a structure
nested deeper than any real database needs, and it is a portable default.
Iterative decoding avoids call-stack exhaustion, but it must still detect cycles
and bound nesting and traversal work.

### Maximum Decoded Value Count

The depth limit does not bound the total amount of work. Arrays and maps can
contain many pointers to the same target. Unless the reader safely reuses the
target, it decodes that target once for every pointer occurrence. In the binary
fan-out example used by the test data, each array contains two pointers to the
level below, so each lower level doubles the work. Decoding the top value takes
exponentially more operations than the depth suggests. A data entry under one
kilobyte can therefore take longer to decode than any real workload allows.

A reader can bound this by counting the values it decodes and stopping if the
count exceeds a fixed limit. One flat accounting rule is:

- Charge every logical value occurrence once. The root is one occurrence. An
array or map is one occurrence in addition to its children. Map keys and map
values are separate occurrences.
- Do not charge a pointer separately from its resolved value. Charge the value
at the position where the pointer occurs. Under this rule, a value resolved
from a memoized target is still charged once for each logical occurrence.

Readers with structural or weighted bounds may account differently.

If a destination uses a container's declared length to allocate storage,
counting only after that allocation may be too late. A reader can reserve the
declared children against its value budget first, cap the allocation separately,
grow storage incrementally, or use another approach with an equivalent bound.

The bound should cover the whole operation the caller requested. It should not
restart between internal phases. For example, path navigation and decoding the
selected value can share one budget or use another end-to-end control. Bounding
each navigation step on its own does not bound their cumulative work.

A limit of 65,536 (2\*\*16) values is recommended. The largest data entries
MaxMind produces decode a few hundred values, so this leaves a wide margin.

### Bounding Expanded Payload

The value-count strategy limits how many values the reader decodes from a data
entry, not how many bytes those values contain. A single string or bytes value
can be up to 16,843,036 bytes. An array of 65,535 pointers to one such value
stays at the recommended value limit under the flat rule above. It can still
describe more than 1 TiB of repeated data in a file barely larger than the value
itself.

A reader that copies, validates, or allocates these values should bound the
total string and bytes data it materializes for one caller-requested decode
operation. The right method and limit depend on the reader's language and API,
so this specification does not require a single limit. As guidance, the largest
data entries MaxMind produces hold about a kilobyte of such data, so a few
megabytes is generous.

Borrowing bytes from the data section or safely memoizing pointer targets can
bound the reader's own materialization. The decoded result may still contain
many logical references to the same value. A binding, serializer, or conversion
to owning values that copies each occurrence should bound that downstream work.

Payload that a reader skips without copying, validating, or allocating need not
consume a materialization budget, as long as the traversal to skip it is bounded
separately. A reader selecting part of a data entry likewise need not expand an
unrequested pointer target. A reader whose API or validation rules require that
work should bound it.

A reader can bound this in several ways:

- Charge each value's size against a per-operation budget wherever it is
decoded, including data stored inline in a container that a pointer targets.
Stop when the total exceeds a limit. Re-decoding a pointer target charges its
data again, which is what bounds the amplification.
- Fold these concerns into a unified weighted budget. It can account for
container entries traversed, pointer expansion when the target is not safely
reused, map-key and selector work, and strings and bytes materialized. A small
fixed-width scalar can cost little or nothing once the work needed to reach it
is bounded.
- Safely memoize decoded pointer targets, handling cycles and in-progress
targets, so a shared target is materialized once.

When a reader's resource controls reject a decode, it should return an error
rather than a partial result. It may treat a value outside its documented
resource profile as invalid, and may make limits configurable for unusually
large valid data entries.

## Reference Implementations

### Writer
Expand Down
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,12 @@ subnets (IPv4 or IPv6).

This repository contains the spec for that format as well as test databases.

Some structurally well-formed databases under `test-data/` are deliberately
hostile and can exhaust an unprotected reader. Do not fully decode every fixture
without resource controls. See the
[denial-of-service test data](test-data/README.md#denial-of-service-test-data)
documentation.

# Generating Test Data

The `write-test-data` command generates the MMDB test files under `test-data/`
Expand Down
5 changes: 5 additions & 0 deletions cmd/write-test-data/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,11 @@ func main() {
os.Exit(1)
}

if err := w.WritePointerDecoderDoSTestDB(); err != nil {
fmt.Printf("writing pointer decoder DoS test databases: %+v\n", err)
os.Exit(1)
}
Comment thread
oschwald marked this conversation as resolved.

if err := w.WriteGeoIP2TestDB(); err != nil {
fmt.Printf("writing GeoIP2 test databases: %+v\n", err)
os.Exit(1)
Expand Down
6 changes: 2 additions & 4 deletions go.mod
Original file line number Diff line number Diff line change
Expand Up @@ -4,10 +4,8 @@ go 1.25.0

require (
github.com/maxmind/mmdbwriter v1.2.0
github.com/oschwald/maxminddb-golang/v2 v2.1.1
go4.org/netipx v0.0.0-20260823151212-3075585bcbeb
)

require (
github.com/oschwald/maxminddb-golang/v2 v2.1.1 // indirect
golang.org/x/sys v0.38.0 // indirect
)
require golang.org/x/sys v0.38.0 // indirect
Loading
Loading