Bound the buffered read to the range the server declared - #39
Conversation
PartialBuffer represents a declared byte range of a remote file, but on the non-stream path it called buffer.read() with no argument, consuming whatever the server chose to send before the declared size was consulted. How much got buffered was decided by the server rather than by the range requested: a client asking for 100 bytes buffered 20 MB when the server streamed that much, which defeats the purpose of fetching ranges at all. Read at most `size` bytes instead, in a loop that tolerates short reads. A single read() call is not enough: a socket-backed response can return fewer bytes than requested while more are still coming, so reading once would silently truncate. The data is written straight into the result buffer so no intermediate copy of the whole range is held. The existing tests all use BytesIO, which never short-reads, so these cases need their own tests. RemoteFetcher.fetch also turned the server's Content-Range straight into that size without checking it. A header whose end precedes its start, such as "bytes 100-50/1000", produced a negative size, and PartialBuffer.read(0) then computed a negative length, which for a file object means read everything. Malformed values such as "bytes abc-def/1000", "bytes */1000" or an empty header raised a bare ValueError out of the library. Both are now RemoteZipError. Adds tests that a server sending more than it declared does not enlarge the buffer, that a server sending less does not hang or raise, that short reads are handled without truncation, and that invalid or malformed Content-Range values are rejected.
gtsystem
left a comment
There was a problem hiding this comment.
Hi, thanks for contributing.
I added one comment. If you can also bump the minor version would be great.
| self.buffer = buffer if stream else io.BytesIO(buffer.read()) | ||
| # Read at most `size` bytes: the declared range is what this buffer | ||
| # represents, and a server may send more than it announced. | ||
| self.buffer = buffer if stream else self._read_up_to(buffer, size) |
There was a problem hiding this comment.
_read_up_to() now stops after size bytes, so the original HTTP response may remain partially consumed. Since buffer is discarded here, its connection is never explicitly closed or released. This can exhaust the HTTP connection pool when a server sends more bytes than declared. Please close the source in a finally block after copying, including calling release_conn() when available.
Example:
if stream:
self.buffer = buffer
else:
try:
self.buffer = self._read_up_to(buffer, size)
finally:
buffer.close()
if hasattr(buffer, 'release_conn'):
buffer.release_conn() # release urllib3 connection associated with this buffer|
Thanks, you're right. Stopping after the declared range left the original response partially consumed, so I updated the non-streaming path to close it and release its connection in a |
|
Released as v0.12.6. thanks again |
PartialBufferrepresents a declared byte range of a remote file, but on thenon-stream path it reads the whole response body:
buffer.read()with no argument consumes whatever the server chooses to send,and it happens before
sizeis consulted at all. So how much gets buffered isdecided by the server rather than by the range requested. Asking for a 100 byte
range from a server that streams 20 MB buffers all 20 MB:
That defeats the point of using range requests, and for anything that runs
remotezip against a URL it does not control it is a way to exhaust memory
remotely.
There is a second route to the same place.
fetchturns the server'sContent-Rangestraight into the size without checking it:A header whose end precedes its start gives a negative size:
PartialBuffer.read(0)then computes a negative length, andfile.read(negative)means read everything for a Python file object. End to end, a client that asked
for 100 bytes received the entire 5000 byte body.
Separately, a header that cannot be parsed leaks a bare
ValueErrorout of thelibrary:
'bytes abc-def/1000','bytes -500-100/1000','bytes /1000'and anempty value all do this today. So does
'bytes */1000', which is theRFC 7233 unsatisfied-range form, so this is not only about hostile input: a
server can send that legitimately.
The change
sizebytes inPartialBuffer, in a loop that tolerates shortreads. A single
read(size)call is not sufficient: a socket-backedresponse can return fewer bytes than requested while more are still coming, so
reading once would silently truncate. Worth stating because the existing tests
all use
BytesIO, which never short-reads, so that failure mode is invisibleto them. The data is written straight into the result buffer, so no
intermediate copy of the whole range is held.
RemoteZipErrorfor aContent-Rangethat ends before it starts orcarries no end, and for one that cannot be parsed at all.
copy, including when copying raises. Rejected
Content-Rangeresponses getthe same cleanup. The streaming path keeps ownership until
PartialBuffer.close()as before.The suffix form that
parse_range_headerreturns forbytes -123is leftalone, since that is a request form rather than a response header, and its
existing test still passes.
bytes 0-99/*, an unknown total length, remainsaccepted and has a test to keep it that way.
What this does not fix
Worth being explicit: this bounds buffering by the size the server declares,
not by the size the client requested. A server that declares a very large
range and then streams it will still be buffered in full:
bytes 0-99/1000bytes 0-4999999/5000000Closing that would mean rejecting or clamping a response range that does not
match the requested one. I have not done it here because the right policy is a
judgement call: some servers legitimately return a different range than asked
for, and
RemoteIO.seekderives_file_sizefrom the declared size, soclamping naively would desynchronise it. Happy to follow up if you have a
preference.
Tests
Six added. Three fail against the current release:
Content-Rangeis rejected;The first two also carry cleanup assertions that fail against the previous PR
head, where their responses remain open. The other three guard adjacent
behaviour:
PartialBuffer.close().Validation
python test_remotezip.py -vgoes from 19 tests to 25, all passing, with noexisting test modified.