ReplayBuffer#
- class torchrl.data.ReplayBuffer(*args, use_ray_service=False, service_backend=None, service_backend_options=None, **kwargs)#
A generic, composable replay buffer class.
See also
ReplayBufferConfig.- Keyword Arguments:
storage (Storage, Callable[[], Storage], optional) – the storage to be used. If a callable is passed, it is used as constructor for the storage. If none is provided a default
ListStoragewithmax_sizeof1_000will be created.sampler (Sampler, Callable[[], Sampler], optional) – the sampler to be used. If a callable is passed, it is used as constructor for the sampler. If none is provided, a default
RandomSamplerwill be used.sample_unit (SampleUnit, optional) – expands the anchors selected by the sampler into the records of the batch (see
SampleUnit).None(default) is equivalent toTransition: every anchor is one transition and classic behavior is preserved.writer (Writer, Callable[[], Writer], optional) – the writer to be used. If a callable is passed, it is used as constructor for the writer. If none is provided a default
RoundRobinWriterwill be used.collate_fn (callable, optional) – merges a list of samples to form a mini-batch of Tensor(s)/outputs. Used when using batched loading from a map-style dataset. The default value will be decided based on the storage type.
pin_memory (bool) – whether pin_memory() should be called on the rb samples.
prefetch (int, optional) – number of next batches to be prefetched using multithreading. Defaults to None (no prefetching).
transform (Transform or Callable[[Any], Any], optional) – Transform to be executed when
sample()is called. To chain transforms use theComposeclass. Transforms should be used withtensordict.TensorDictcontent. A generic callable can also be passed if the replay buffer is used with PyTree structures (see example below). Unlike storages, writers and samplers, transform constructors must be passed as separate keyword argumenttransform_factory, as it is impossible to distinguish a constructor from a transform.transform_factory (Callable[[], Callable], optional) – a factory for the transform. Exclusive with
transform.batch_size (int, optional) –
the batch size to be used when sample() is called.
Note
The batch-size can be specified at construction time via the
batch_sizeargument, or at sampling time. The former should be preferred whenever the batch-size is consistent across the experiment. If the batch-size is likely to change, it can be passed to thesample()method. This option is incompatible with prefetching (since this requires to know the batch-size in advance) as well as with samplers that have adrop_lastargument.dim_extend (int, optional) –
indicates the dim to consider for extension when calling
extend(). Defaults tostorage.ndim-1. When usingdim_extend > 0, we recommend using thendimargument in the storage instantiation if that argument is available, to let storages know that the data is multi-dimensional and keep consistent notions of storage-capacity and batch-size during sampling.Important
When using a collector with
trajs_per_batch, trajectories are written as flat 1-D sequences of variable length. Do not setdim_extend > 0orndim >= 2in this case — the storage must be 1-dimensional.Note
This argument has no effect on
add()and therefore should be used with caution when bothadd()andextend()are used in a codebase. For example:>>> data = torch.zeros(3, 4) >>> rb = ReplayBuffer( ... storage=LazyTensorStorage(10, ndim=2), ... dim_extend=1) >>> # these two approaches are equivalent: >>> for d in data.unbind(1): ... rb.add(d) >>> rb.extend(data)
generator (torch.Generator, optional) –
a generator to use for sampling. Using a dedicated generator for the replay buffer can allow a fine-grained control over seeding, for instance keeping the global seed different but the RB seed identical for distributed jobs. Defaults to
None(global default generator).Warning
As of now, the generator has no effect on the transforms.
consume_after_n_samples (int, optional) – if provided, sampled items are removed from the sampleable set after they have been returned this many times. The default value of
Nonekeeps the standard replay buffer behavior. Passing1makes each item available for a single sample before it is consumed.shared (bool, optional) – whether the buffer will be shared using multiprocessing or not. Defaults to
False.compilable (bool, optional) – whether the writer is compilable. If
True, the writer cannot be shared between multiple processes. Defaults toFalse.delayed_init (bool, optional) – whether to initialize storage, writer, sampler and transform the first time the buffer is used rather than during construction. This is useful when the replay buffer needs to be pickled and sent to remote workers, particularly when using transforms with modules that require gradients. If not specified, defaults to
Truewhentransform_factoryis provided, andFalseotherwise.service_backend (str) – deployment backend, either
"direct"or"ray". Defaults to"direct".service_backend_options (dict, optional) – Ray initialization options. Accepted keys are
ray_init_configandremote_config.transport (str, optional) – physical transport used by a remote replay owner.
"auto"selects the backend default. Defaults to"auto".transport_options (dict, optional) – options for the selected transport. For
transport="distributed",backendselects"gloo"or"nccl". TensorDict layouts are bound lazily on first use.
Examples
>>> import torch >>> >>> from torchrl.data import ReplayBuffer, ListStorage >>> >>> torch.manual_seed(0) >>> rb = ReplayBuffer( ... storage=ListStorage(max_size=1000), ... batch_size=5, ... ) >>> # populate the replay buffer and get the item indices >>> data = range(10) >>> indices = rb.extend(data) >>> # sample will return as many elements as specified in the constructor >>> sample = rb.sample() >>> print(sample) tensor([4, 9, 3, 0, 3]) >>> # Passing the batch-size to the sample method overrides the one in the constructor >>> sample = rb.sample(batch_size=3) >>> print(sample) tensor([9, 7, 3]) >>> # one cans sample using the ``sample`` method or iterate over the buffer >>> for i, batch in enumerate(rb): ... print(i, batch) ... if i == 3: ... break 0 tensor([7, 3, 1, 6, 6]) 1 tensor([9, 8, 6, 6, 8]) 2 tensor([4, 3, 6, 9, 1]) 3 tensor([4, 4, 1, 9, 9])
Replay buffers accept any kind of data. Not all storage types will work, as some expect numerical data only, but the default
ListStoragewill:Examples
>>> torch.manual_seed(0) >>> buffer = ReplayBuffer(storage=ListStorage(100), collate_fn=lambda x: x) >>> indices = buffer.extend(["a", 1, None]) >>> buffer.sample(3) [None, 'a', None]
The
TensorStorage,LazyMemmapStorageandLazyTensorStoragealso work with any PyTree structure (a PyTree is a nested structure of arbitrary depth made of dicts, lists or tuples where the leaves are tensors) provided that it only contains tensor data.Examples
>>> from torch.utils._pytree import tree_map >>> def transform(x): ... # Zeros all the data in the pytree ... return tree_map(lambda y: y * 0, x) >>> rb = ReplayBuffer(storage=LazyMemmapStorage(100), transform=transform) >>> data = { ... "a": torch.randn(3), ... "b": {"c": (torch.zeros(2), [torch.ones(1)])}, ... 30: -torch.ones(()), ... } >>> rb.add(data) >>> # The sample has a similar structure to the data (with a leading dimension of 10 for each tensor) >>> s = rb.sample(10) >>> # let's check that our transform did its job: >>> def assert0(x): >>> assert (x == 0).all() >>> tree_map(assert0, s)
- add(data: Any) int[source]#
Add a single element to the replay buffer.
- Parameters:
data (Any) – data to be added to the replay buffer
- Returns:
index where the data lives in the replay buffer.
- append_transform(transform: Transform, *, invert: bool = False) ReplayBuffer[source]#
Appends transform at the end.
Transforms are applied in order when sample is called.
- Parameters:
transform (Transform) – The transform to be appended
- Keyword Arguments:
invert (bool, optional) – if
True, the transform will be inverted (forward calls will be called during writing and inverse calls during reading). Defaults toFalse.
Example
>>> rb = ReplayBuffer(storage=LazyMemmapStorage(10), batch_size=4) >>> data = TensorDict({"a": torch.zeros(10)}, [10]) >>> def t(data): ... data += 1 ... return data >>> rb.append_transform(t, invert=True) >>> rb.extend(data) >>> assert (data == 1).all()
- classmethod as_remote(remote_config=None)#
Creates an instance of a remote ray class.
- Parameters:
cls (Python Class) – class to be remotely instantiated.
remote_config (dict) – the quantity of CPU cores to reserve for this class. Defaults to torchrl.collectors.distributed.ray.DEFAULT_REMOTE_CLASS_CONFIG.
- Returns:
A function that creates ray remote class instances.
- property batch_size#
The batch size of the replay buffer.
The batch size can be overridden by setting the batch_size parameter in the
sample()method.It defines both the number of samples returned by
sample()and the number of samples that are yielded by theReplayBufferiterator.
- dumps(path)[source]#
Saves the replay buffer on disk at the specified path.
- Parameters:
path (Path or str) – path where to save the replay buffer.
Examples
>>> import tempfile >>> import tqdm >>> from torchrl.data import LazyMemmapStorage, TensorDictReplayBuffer >>> from torchrl.data.replay_buffers.samplers import PrioritizedSampler, RandomSampler >>> import torch >>> from tensordict import TensorDict >>> # Build and populate the replay buffer >>> S = 1_000_000 >>> sampler = PrioritizedSampler(S, 1.1, 1.0) >>> # sampler = RandomSampler() >>> storage = LazyMemmapStorage(S) >>> rb = TensorDictReplayBuffer(storage=storage, sampler=sampler) >>> >>> for _ in tqdm.tqdm(range(100)): ... td = TensorDict({"obs": torch.randn(100, 3, 4), "next": {"obs": torch.randn(100, 3, 4)}, "td_error": torch.rand(100)}, [100]) ... rb.extend(td) ... sample = rb.sample(32) ... rb.update_tensordict_priority(sample) >>> # save and load the buffer >>> with tempfile.TemporaryDirectory() as tmpdir: ... rb.dumps(tmpdir) ... ... sampler = PrioritizedSampler(S, 1.1, 1.0) ... # sampler = RandomSampler() ... storage = LazyMemmapStorage(S) ... rb_load = TensorDictReplayBuffer(storage=storage, sampler=sampler) ... rb_load.loads(tmpdir) ... assert len(rb) == len(rb_load)
- empty(empty_write_count: bool = True)[source]#
Empties the replay buffer and reset cursor to 0.
- Parameters:
empty_write_count (bool, optional) – Whether to empty the write_count attribute. Defaults to True.
- extend(data: Sequence, *, update_priority: bool | None = None) Tensor[source]#
Extends the replay buffer with one or more elements contained in an iterable.
If present, the inverse transforms will be called.`
- Parameters:
data (iterable) – collection of data to be added to the replay buffer.
- Keyword Arguments:
update_priority (bool, optional) – Whether to update the priority of the data. Defaults to True. Without effect in this class. See
extend()for more details.- Returns:
Indices of the data added to the replay buffer.
Warning
extend()can have an ambiguous signature when dealing with lists of values, which should be interpreted either as PyTree (in which case all elements in the list will be put in a slice in the stored PyTree in the storage) or a list of values to add one at a time. To solve this, TorchRL makes the clear-cut distinction between list and tuple: a tuple will be viewed as a PyTree, a list (at the root level) will be interpreted as a stack of values to add one at a time to the buffer. ForListStorageinstances, only unbound elements can be provided (no PyTrees).
- property initialized: bool#
Whether the replay buffer has been initialized.
- insert_transform(index: int, transform: Transform, *, invert: bool = False) ReplayBuffer[source]#
Inserts transform.
Transforms are executed in order when sample is called.
- Parameters:
index (int) – Position to insert the transform.
transform (Transform) – The transform to be appended
- Keyword Arguments:
invert (bool, optional) – if
True, the transform will be inverted (forward calls will be called during writing and inverse calls during reading). Defaults toFalse.
- property is_alive: bool#
Whether this direct replay buffer remains available.
- loads(path)[source]#
Loads a replay buffer state at the given path.
The buffer should have matching components and be saved using
dumps().- Parameters:
path (Path or str) – path where the replay buffer was saved.
See
dumps()for more info.
- next()[source]#
Returns the next item in the replay buffer.
This method is used to iterate over the replay buffer in contexts where __iter__ is not available, such as
RayReplayBuffer.
- query(predicate: Callable[[Trajectory], bool] | None = None, *, trajectory_key: NestedKey | None = None) list[Trajectory][source]#
Filters the stored trajectories with a query predicate.
Splits the buffer content into trajectories (see
iter_trajectories()) and returns those matching the predicate asTrajectoryviews.- Parameters:
predicate (Callable[[Trajectory], bool], optional) – a
TrajectoryPredicatebuilt fromtraj, or any callable mapping a trajectory to a boolean. Defaults to None (return all trajectories).- Keyword Arguments:
trajectory_key (NestedKey, optional) – entry holding per-transition trajectory ids. Defaults to None (auto-detection from
("collector", "traj_ids"),"traj_ids","episode"or the done/terminated/truncated flags).- Returns:
A list of matching trajectory views, ordered chronologically (oldest trajectory first; for multi-dimensional storages, grouped by batch coordinate).
The trajectory boundaries are computed from the stored (untransformed) data with the same machinery
SliceSampleruses, so samplers and queries always agree on where trajectories start and stop. This includes storages withndim > 1(e.g.LazyTensorStorage(..., ndim=2)holding[B, T]batches), whose trajectories are recovered per batch coordinate.Predicates built from
trajreport the keys they read viarequired_keys(); evaluation then only fetches those entries from the storage and only runs the transforms that can affect them. Matching trajectories are extracted in full with the complete transform chain applied, so predicates and results see the same values a sampler would produce. Opaque callables are evaluated against the fully transformed content.Note
Once the buffer has wrapped around (it is at capacity and older entries have been overwritten), the oldest trajectory may have lost its first transitions to overwriting and will appear truncated at the front. A trajectory written across the wrap point is followed through it and returned whole, in time order.
Examples
>>> from torchrl.data import traj >>> good_trajs = rb.query((traj.reward.sum() > 100) & (traj.length >= 50)) >>> observations = good_trajs[0].observation
- read_all_in_order(end: int | None = None) Any[source]#
Read storage contents in physical order.
This is equivalent to
rb[:]whenendisNone.- Parameters:
end (int, optional) – Number of leading storage entries to read. Defaults to the entire storage slice.
- Returns:
A storage slice containing entries
[:end].
- register_load_hook(hook: Callable[[Any], Any])[source]#
Registers a load hook for the storage.
Note
Hooks are currently not serialized when saving a replay buffer: they must be manually re-initialized every time the buffer is created.
- register_save_hook(hook: Callable[[Any], Any])[source]#
Registers a save hook for the storage.
Note
Hooks are currently not serialized when saving a replay buffer: they must be manually re-initialized every time the buffer is created.
- sample(batch_size: int | None = None, return_info: bool = False) Any[source]#
Samples a batch of data from the replay buffer.
Uses Sampler to sample indices, and retrieves them from Storage.
- Parameters:
batch_size (int, optional) – size of data to be collected. If none is provided, this method will sample a batch-size as indicated by the sampler.
return_info (bool) – whether to return info. If True, the result is a tuple (data, info). If False, the result is the data.
- Returns:
A batch of data selected in the replay buffer. A tuple containing this batch and info if return_info flag is set to True.
- property sampler: Sampler#
The sampler of the replay buffer.
The sampler must be an instance of
Sampler.
- property service_backend: str#
The canonical deployment backend for this replay buffer.
- set_(key, value)[source]#
Sets the value of a key across the entire replay buffer in-place.
- Parameters:
key (NestedKey) – the key to set.
value (torch.Tensor) – the value to write.
- Returns:
self
- set_at_(key, value, index)[source]#
Sets the value of a key at specified indices in the replay buffer.
- Parameters:
key (NestedKey) – the key to set.
value (torch.Tensor) – the value to write.
index – the indices where to write the value.
- Returns:
self
- set_sampler(sampler: Sampler)[source]#
Sets a new sampler in the replay buffer and returns the previous sampler.
- set_storage(storage: Storage, collate_fn: Callable | None = None)[source]#
Sets a new storage in the replay buffer and returns the previous storage.
- Parameters:
storage (Storage) – the new storage for the buffer.
collate_fn (callable, optional) – if provided, the collate_fn is set to this value. Otherwise it is reset to a default value.
- set_writer(writer: Writer)[source]#
Sets a new writer in the replay buffer and returns the previous writer.
- shutdown(timeout: float | None = None) None[source]#
Mark this direct replay-buffer owner as shut down.
- stats() dict[str, int | float | bool][source]#
Returns a cheap, serializable snapshot of the buffer’s operational state.
The snapshot only contains scalar counters and gauges. It never includes the storage content, does not modify the buffer state and is safe to call concurrently with writes and samples. Cumulative counters such as
write_countare meant to be converted into rates by an external monitor such asLoggerMonitor.Calling this method on an uninitialized buffer does not trigger its initialization; an empty snapshot with
initialized=Falseis returned instead (capacityis still reported when the storage already advertises it).- Returns:
"size": current number of elements in the buffer (mirrorslen(buffer));"write_count": total number of items written throughaddandextend(0for writers that do not track writes, such asImmutableDatasetWriter);"prefetch_queue_size": number of pending prefetched batches;"initialized": whether the buffer components are initialized;"capacity": maximum number of elements the storage can hold (only present when the storage advertises amax_size);"utilization":size / capacity(only present alongsidecapacity).
Remote clients backed by the distributed transport report a subset of these entries (
sizeandwrite_count).- Return type:
A dictionary with the following entries
Examples
>>> import torch >>> from torchrl.data import LazyTensorStorage, ReplayBuffer >>> rb = ReplayBuffer(storage=LazyTensorStorage(10)) >>> rb.extend(torch.arange(5)) >>> snapshot = rb.stats() >>> print(snapshot["size"], snapshot["write_count"], snapshot["capacity"]) 5 5 10
- property storage: Storage#
The storage of the replay buffer.
The storage must be an instance of
Storage.
- property transform: Transform#
The transform of the replay buffer.
The transform must be an instance of
Transform.
- update_(input_dict_or_td, clone=False, *, keys_to_update=None)[source]#
Updates the replay buffer in-place with the given dict or TensorDict.
- Parameters:
input_dict_or_td (dict or TensorDictBase) – the data to update with.
clone (bool, optional) – whether to clone the values before writing. Defaults to
False.keys_to_update (sequence of NestedKey, optional) – if provided, only these keys will be updated.
- Returns:
self
- update_if_present(*, index: Tensor, generation: Tensor, patch: Mapping[NestedKey, Tensor] | TensorDictBase, version_key: NestedKey | None = None, version: int | Tensor | None = None, require_newer: bool = False) ConditionalUpdateResult[source]#
Conditionally updates stored records that are still live.
Replay slots are recycled by round-robin writers, so a physical index captured at sampling time can point to a different record by the time an asynchronous computation writes back. This method applies
patchonly to records whose(index, generation)pair still matches the writer’s current slot generation, skipping records whose slot was reused or emptied since the handle was captured. Skipped records are never modified.The whole patch is validated (key existence, shape and dtype) before any write happens; a validation failure leaves the storage untouched. Updating a record refreshes its content, not its identity: the same handle keeps working until the slot is rewritten by
add,extendorempty.Generation tracking is opt-in: the buffer must be constructed with a writer that tracks slot generations, e.g.
RoundRobinWriter(track_generations=True)(see ref_buffers_generations). Calling this method on a buffer whose writer does not track generations raises aRuntimeError.- Keyword Arguments:
index (torch.Tensor) – storage indices, as returned by
extend()or found in the sample under"index".generation (torch.Tensor) – slot generations captured with the indices, as found in the sample under
"index_generation".patch (mapping of NestedKey to torch.Tensor, or TensorDictBase) – the fields to overwrite for live records. Leading dimension must match the number of records addressed by
index.version_key (NestedKey, optional) – a stored per-record scalar field holding each record’s current version. When passed (together with
version), a generation-live record is only patched if the incoming version compares favorably against the stored one, and the accepted version is written intoversion_keyatomically with the patch.version_keymay not appear inpatch. Nested keys must be passed in tuple form (("nested", "version")); dotted strings are rejected. Defaults toNone(no version comparison).version (int or torch.Tensor, optional) – the incoming version, either a scalar (broadcast to every record) or a tensor with one entry per record. Must be passed together with
version_key.require_newer (bool, optional) – if
True, a record is only patched whenversion > stored; ifFalse, ties are accepted (version >= stored). When the same slot is addressed several times in one call, only the row carrying the highest incoming version is applied (the last such row on ties); the losing rows are reported inversion_rejected. Defaults toFalse.
- Returns:
A
ConditionalUpdateResultwhoseupdatedmask is aligned with the input index order, withupdated_countandstale_countconveniences. Whenversion_keyis passed, itsversion_rejectedmask marks generation-live records that were rejected by the version comparison (Noneotherwise).- Raises:
RuntimeError – if the storage does not support conditional updates (for example
ListStorage) or the writer does not track slot generations.KeyError – if a patch key (or
version_key) does not exist in the storage.ValueError – if a patch entry has an incompatible shape or dtype, if only one of
version_key/versionis passed, ifversion_keyappears inpatchor names a non-scalar field, or if it is a dotted string.
Examples
>>> import torch >>> from tensordict import TensorDict >>> from torchrl.data import ( ... LazyTensorStorage, ... TensorDictReplayBuffer, ... TensorDictRoundRobinWriter, ... ) >>> rb = TensorDictReplayBuffer( ... storage=LazyTensorStorage(10), ... writer=TensorDictRoundRobinWriter(track_generations=True), ... batch_size=4, ... ) >>> rb.extend(TensorDict({"obs": torch.zeros(10, 3)}, batch_size=[10])) >>> sample = rb.sample() >>> result = rb.update_if_present( ... index=sample["index"], ... generation=sample["index_generation"], ... patch={"obs": torch.ones(4, 3)}, ... ) >>> print(result.updated_count, result.stale_count) 4 0
With a version comparison, outdated asynchronous writers lose deterministically:
>>> rb = TensorDictReplayBuffer( ... storage=LazyTensorStorage(10), ... writer=TensorDictRoundRobinWriter(track_generations=True), ... batch_size=4, ... ) >>> rb.extend( ... TensorDict( ... { ... "obs": torch.zeros(10, 3), ... "v": torch.full((10,), 5, dtype=torch.int64), ... }, ... batch_size=[10], ... ) ... ) >>> sample = rb.sample() >>> result = rb.update_if_present( ... index=sample["index"], ... generation=sample["index_generation"], ... patch={"obs": torch.ones(4, 3)}, ... version_key="v", ... version=4, ... require_newer=True, ... ) >>> print(result.updated_count, result.version_rejected_count) 0 4
- write_all(data: Any, end: int | None = None) None[source]#
Write data back to storage in physical order.
This is equivalent to
rb[:end] = data. IfendisNone,enddefaults todata.shape[0]for tensor collections andlen(data)otherwise. Ifdataspans the full storage, this is equivalent torb[:] = data.- Parameters:
data – Data to write to storage.
end (int, optional) – Number of leading storage entries to update. Defaults to
data.shape[0]for tensor collections andlen(data)otherwise.
- property write_count: int#
The total number of items written so far in the buffer through add and extend.