Repository navigation
chacha20: 64-bit counter support #334
Description
Activity
I have some seemingly working code for this, but it builds on the Rng PR with
variants.rs. Would you like me to try to adjust it for the current state ofchacha20or wait for the potential merge before opening a PR for it?Let's try to land #333 first
I took another look at my 64-bit counter code, and it isn't working correctly on the NEON backend atm.
It uses a const bool in
variants.rsto determine which way the counter should be added in the backend, but if there's a better way to do this that doesn't involve modifying the backend code... it might be better. It is currently failing a test shortly after reaching the u32::MAX block pos.The largest use-case for a 64-bit counter (that I can think of) is file system encryption for mobile devices. I've ran some benches, and
chacha20and evenchacha20poly1305is faster than AES onaarch64. I've benched it on the cloud, on a compute EC2 instance. I would imagine that it's also faster on mobile devices. So... if the 64-bit counter only works for x86 and the software implementation, then there wouldn't really be a point if the main use case is on aarch64.I will keep working on it a little, but I'd probably prefer a solution that doesn't need to modify the backends, if that's possible. Because the way it is now, every new backend would also need to have specific code for updating the counter
On second thought, I feel like more control over a larger nonce would be favorable over having a larger counter
@nstilt1 not sure what you mean by that? Are you suggesting making the construction generic over the nonce size rather than a larger counter?
@tarcieri I'm suggesting that 64-bit counters are unnecessary, contrary to what I've said previously. I think having 96 bits for a nonce is more practical because it is easier to prevent nonce reuse with finer control over the value. What do you mean about making the construction generic over the nonce size?
I wasn't sure what you meant by "more control over a larger nonce".
I agree practically larger counter support probably doesn't matter which is why the current implementation avoided dealing with it. I'm not sure the added complexity is worth it, though it could be considered a compatibility bug with the legacy implementation which is somewhat surprisingly pervasive.
Here is my use case which currently panics, but would not panic (I think) with 64 bit counters:
Encrypting a stream of files to backup a file to S3 object. It is reasonable that I might have a stream that is >256GiB long.
I did not write the backup program yet. It is in progress. Here is something I tried out if I imagine continuing an in progress upload at 500GiB:
let key = [0x42; 32]; let nonce = [0x24; 12]; let mut cipher = ChaCha20::new(&key.into(), &nonce.into()); const SEEK_AMOUNT: usize = 536_870_912_000; cipher.seek(SEEK_AMOUNT);
But currently the
seekpanics.I am new to file encryption so let me know if my use case is invalid and there is something else I should be using other than ChaCha20 (I don't need protection against the data being modified, I just want the files to not be readable without the key).
You could change the nonce every so often, such as increment a portion of the nonce once every 256 GiB of encrypted data. But I could also revisit #359. I'll probably use a different PR since that one is kinda far behind.
The only problem is that the code would be slightly different if #380 were to be merged. I could put the counter support in #380, or #380 could just be closed and I could put the counter support in a new PR. Also regarding this:
I'm not sure the added complexity is worth it
The changes made in #359 worked. If I'm going to revisit this issue, the backends will look almost exactly like they do in that PR. If desired, I could make some adjustments, such as making an
add_counter!macro for each SIMD backend. That makes it so that the counter logic is written once per backend, reducing the chances of a typo while (supposedly) maintaining the compile-time evaluation of the logic. The macro would look similar toadd_counter!inneon.rs, and I personally think that using this macro reads a little bit better than without the macro.@ChocolateLoverRaj are you worried about encrypting a single file that is 256 GiB, or the sum total of multiple files being larger?
If it's the latter, you should split up the encryption so each file is encrypted under a different key/nonce.
@ChocolateLoverRaj are you worried about encrypting a single file that is 256 GiB, or the sum total of multiple files being larger?
My plan is to create a stream of all the file changes (including the contents of all new files). So every file will be <256GiB but all of the files together might result in a backup >256GiB. I will be splitting the stream up into 5GB chunks (because AWS limit), but I would prefer to have a single cipher for the entire big stream and split up the encrypted stream rather than have to create multiple ciphers for different chunks of the stream.
If I do split it up the stream and use multiple ciphers, can I use the same key and increment the nonce (like using a nonce of
0for part 0 and1for part 1, etc), as long as I don't use the same nonce twice?If I do split it up the stream and use multiple ciphers, can I use the same key and increment the nonce (like using a nonce of 0 for part 0 and 1 for part 1, etc), as long as I don't use the same nonce twice?
Yes though for large files it would be better to use a unique key per file
Yes though for large files it would be better to use a unique key per file
Why?
256 GB is the data volume limit of e.g. Poly1305 as an authenticator, which you should be using in tandem with ChaCha20 in the combined ChaCha20Poly1305 AEAD cipher to prevent chosen ciphertext attacks
256GB is also generally quite a bit beyond what most computers can hold in RAM, and really you should be working with authenticated AEAD messages you can hold in RAM.
6 remaining items
- marked ChaCha counter is limited to 32-bit block-ciphers#488 as a duplicate of this issue
on Jun 25, 2025 It seems 64-bit support is the main blocker for replacing
rand_chachawithchacha20: rust-random/rand#1642 (review)Indeed, 32-bit counters are insufficient for general-purpose PRNGs. 64-bit is sufficient given that ChaCha supports independent streams.
I have some time if you want me to implement it. Here are some options, each one should require approximately 2 commits, one for x86/software and one for aarch64.
- Use a generic/type arg to determine the counter. Possibly make StreamID and block pos for both types, eg:
type StreamID = LegacyStreamID - Use 64-bit counter solely in the backends, and only use 64-bit counter for RNG. Gatekeep the 64-bit functionality for
ChaCha20and panic when 32 bit counter is exhausted
There are downsides to sticking with only one counter, such as only the 64-bit counter. To be specific, I would need to refactor my personal crypto protocol that uses the RNG from
chacha20. It would be easy to refactor, but it would require me to refactor my library nonetheless so that the proper words in the state are set to the correct values. For those who are curious, I would specifically need to refactor this so that the 13th word in the state is set to the correct value. I know I kind of wasted a word by duplicating it, and could have completely avoided having to refactor by just putting in0or a constant there, but I kind of wanted to use all 3 words in the stream ID. I suppose I could have used variables forsaltorkeyas the first word, what it's for for the second word (eccormacorsymmetric keyorrsa), and theversionfor the third word.- Use a generic/type arg to determine the counter. Possibly make StreamID and block pos for both types, eg:
@nstilt1 the other option I was suggesting was having the backends use a 32-bit counter internally, and implement newtype which provides a 64-bit counter by incrementing the lower 32-bit portion of the nonce when the 32-bit cipher has exhausted its keystream
@tarcieri Would that be a wrapper? Or a different
$ChaChaXRngLegacysort of thing with a different macro to implement it? Or implemented in the same macro that the current one is implemented in? I feel like we would need to restrict the internal block pos to be a multiple of 4 if it is not already. Or would this "wrapper" be implemented below the$ChaChaXRngand above the core, or replacing the core?I will do some poking around tonight. If it just requires changes to
rng.rsI might be able to figure it out. I'm thinking of making the macro impl both ChaChaXRng as well as ChaChaXLegacyRng because I don't want to have breaking changes. Might be possible to just create an alternate rng with the macro.How should the
StreamIDfunction? A 64-bit stream ID and a 96-bit stream ID, or just one stream ID type that is locked at 64 bits or 96 bits, or just one stream ID type that can be constructed with either 64 bits or 96 bits? I'm leaning towards two stream ID types since the behavior could get kind of wacky after the counter surpasses 32 bits, potentially causing bugs.@nstilt1 do you not like the direction of #399?
I don't think 64-bit counters can be limited to
rng.rssince backends generate multiple blocks at once which need to increment the counter (and might even wrap here: no one enforces that the counter is a multiple of 4).In my opinion
StreamIdshould just beu64.Would that be a wrapper? Or a different $ChaChaXRngLegacy sort of thing
@nstilt1 one way it could be implemented is as a newtype of
ChaChaCore(i.e. astructwith an innerChaChaCoreinstance)The main thing I think it would be nice to avoid is having to make every backend generic over both a 32-bit and 64-bit counter.
Oh. I did not see 399. And I now think I may have been wrong about the multiple of four thing. Maybe if the backend was using a 64-bit counter it would have been true, I think I was still in the headspace of using only 64-bit counters in the backend. I'll finish tinkering with
rng.rsin a few moments and will check out that prI see that they did not modify
rng.rs. And the lowest I can go inrng.rsis the generate function, which could be used to increment the counter when it needs to, but it works better if I change the set word pos and block pos functions to set the counter to a multiple of 4 and adjust the index to retain the desired word pos. How should I go about adding 64-bit counter support for the RNG? I've got a PR here: #439I've successfully implemented it with all tests passing, but there aren't any tests covering what happens after the counter carries over, and it's missing the 64-bit getters and setters still. Added a task list to the pr. If you would rather stick with the other PR, I can close mine. I only spent one evening on it after all
It is done now. It satisfies everyone's requests: 64-bit counter for the RNG, 32-bit counter functionality for the IETF
ChaCha20cipher, no generics for the counter code in the backends. If you can write a test that fails by getting the cipher's counter to overflow, go ahead and do it so we can fix it. I tried it with no luck. Might only be possible to overflow the counter usingChaChaCore- added a commit that references this issue
on Aug 27, 2025
The
ChaCha20Legacyconstruction, i.e. the djb variant, is supposed to use a 64-bit counter but currently uses a 32-bit counter because it shares its core implementation with the IETF construction which uses a 32-bit counter.This results in a counter overflow after generating 256 GiB of keystream. Compatible implementations are able to generate larger keystreams.
I'm not sure how much of a practical concern this actually is, but it did come up in discussions here: rust-random/rand#934 (comment)
We can probably make the counter type generic between
u32/u64in the core implementation if need be.