InplaceFunction#
- class torch.autograd.function.InplaceFunction(inplace=False)[source]#
This class is here only for backward compatibility reasons. Use
Functioninstead of this for any new use case.- static backward(ctx, *grad_outputs)[source]#
Define a formula for differentiating the operation with backward mode automatic differentiation.
This function is to be overridden by all subclasses. (Defining this function is equivalent to defining the
vjpfunction.)It must accept a context
ctxas the first argument, followed by as many outputs as theforward()returned (None will be passed in for non tensor outputs of the forward function), and it should return as many tensors, as there were inputs toforward(). Each argument is the gradient w.r.t the given output, and each returned value should be the gradient w.r.t. the corresponding input. If an input is not a Tensor or is a Tensor not requiring grads, you can just pass None as a gradient for that input.The strides of the gradients passed to
backward()are undefined: they are not guaranteed to be contiguous or to match the strides of the corresponding forward outputs, so implementations must not assume a particular memory layout.The context can be used to retrieve tensors saved during the forward pass. It also has an attribute
ctx.needs_input_gradas a tuple of booleans representing whether each input needs gradient. E.g.,backward()will havectx.needs_input_grad[0] = Trueif the first input toforward()needs gradient computed w.r.t. the output.- Return type:
- static forward(*args, **kwargs)[source]#
Define the forward of the custom autograd Function.
This function is to be overridden by all subclasses. There are two ways to define forward:
Usage 1 (Combined forward and ctx):
@staticmethod def forward(ctx: Any, *args: Any, **kwargs: Any) -> Any: pass
It must accept a context ctx as the first argument, followed by any number of arguments (tensors or other types).
See Combined or separate forward() and setup_context() for more details
Usage 2 (Separate forward and ctx):
@staticmethod def forward(*args: Any, **kwargs: Any) -> Any: pass @staticmethod def setup_context(ctx: Any, inputs: Tuple[Any, ...], output: Any) -> None: pass
The forward no longer accepts a ctx argument.
Instead, you must also override the
torch.autograd.Function.setup_context()staticmethod to handle setting up thectxobject.outputis the output of the forward,inputsare a Tuple of inputs to the forward.See Extending torch.autograd for more details
The context can be used to store arbitrary data that can be then retrieved during the backward pass. Tensors should not be stored directly on ctx (though this is not currently enforced for backward compatibility). Instead, tensors should be saved either with
ctx.save_for_backward()if they are intended to be used inbackward(equivalently,vjp) orctx.save_for_forward()if they are intended to be used for injvp.- Return type:
- property input_grad_buffers: tuple[Tensor | None, ...]#
Return existing buffers for accumulating gradients of this Function’s inputs.
Each entry corresponds to an argument passed to
forward(), in the same order. Each entry isNoneor the autograd engine’s currentInputBufferfor that input. A non-Nonebuffer contains gradient contributions already produced during the current backward. A custom backward may accumulate its contribution directly into the buffer and returnNonefor that input, fusing gradient computation with accumulation and avoiding a separate gradient tensor.Availability follows backward execution order and is independent for each input. An entry is
Nonewhen no earlier producer has contributed to that input, or when the existing buffer is aliased or cannot safely be updated in place.For example,
xhas another forward use whileweightdoes not. The custom backward conditionally fuses accumulation only forgrad_x:>>> class Matmul(torch.autograd.Function): >>> @staticmethod >>> def forward(ctx, x, weight): >>> ctx.save_for_backward(x, weight) >>> return x @ weight >>> >>> @staticmethod >>> def backward(ctx, grad_output): >>> x, weight = ctx.saved_tensors >>> x_buffer, _ = ctx.input_grad_buffers >>> if x_buffer is not None: >>> # Computes grad_x and adds it to the existing buffer. >>> matmul_backward_input_acc( >>> grad_output, weight, acc_into=x_buffer >>> ) >>> grad_x = None >>> else: >>> grad_x = matmul_backward_input(grad_output, weight) >>> grad_weight = matmul_backward_weight(grad_output, x) >>> return grad_x, grad_weight >>> >>> loss = Matmul.apply(x, weight).sum() + other_op(x).sum()
If
other_opproduces its contribution first,x_buffercan expose that partial sum. IfMatmulruns first,x_bufferisNoneand it returns a separate tensor instead. The fallback lets the function work under either ordering.Warning
A returned buffer is valid only while the current custom
backwardinvocation is running. Mutate it synchronously and do not retain it. A later producer may replace the engine’s buffer, making a retained tensor stale.After receiving a non-
Nonebuffer, callingbackwardorgradbefore the custom backward returns raises an error.All engine-scheduled producers that use or subsequently update an exposed buffer must execute on the same device, autograd engine thread, and stream.
This property is available only while a Python custom
backwardis executing during an eager, first-orderbackward(),torch.autograd.backward(), ortorch.autograd.grad()call. It is unavailable withcreate_graph=True, anomaly detection, a post-hook on the producing autograd node, or stale capture stream overrides.Note
For a leaf input, a non-
Noneentry exposes its execution-localInputBuffer, not its existing.grad. All contributions are first combined in that buffer. Duringbackward,AccumulateGradthen runs once with the completed buffer to update.gradand run its usual hooks.torch.autograd.grad()instead returns the completed buffer without updating.grad.During
backward, a custom backward that instead accumulates directly into a leaf.gradand returnsNonedoes not use this interface. It is responsible for managing.gradstate, including initialization and lifetime, synchronization with all other producers, and anyAccumulateGradhook behavior bypassed by the direct write.
- static jvp(ctx, *grad_inputs)[source]#
Define a formula for differentiating the operation with forward mode automatic differentiation.
This function is to be overridden by all subclasses. It must accept a context
ctxas the first argument, followed by as many inputs as theforward()got (None will be passed in for non tensor inputs of the forward function), and it should return as many tensors as there were outputs toforward(). Each argument is the gradient w.r.t the given input, and each returned value should be the gradient w.r.t. the corresponding output. If an output is not a Tensor or the function is not differentiable with respect to that output, you can just pass None as a gradient for that input.You can use the
ctxobject to pass any value from the forward to this functions.- Return type:
- mark_dirty(*args)[source]#
Mark given tensors as modified in an in-place operation.
This should be called at most once, in either the
setup_context()orforward()methods, and all arguments should be inputs.Every tensor that’s been modified in-place in a call to
forward()should be given to this function, to ensure correctness of our checks. It doesn’t matter whether the function is called before or after modification.- Examples::
>>> class Inplace(Function): >>> @staticmethod >>> def forward(ctx, x): >>> x_npy = x.numpy() # x_npy shares storage with x >>> x_npy += 1 >>> ctx.mark_dirty(x) >>> return x >>> >>> @staticmethod >>> @once_differentiable >>> def backward(ctx, grad_output): >>> return grad_output >>> >>> a = torch.tensor(1., requires_grad=True, dtype=torch.double).clone() >>> b = a * a >>> Inplace.apply(a) # This would lead to wrong gradients! >>> # but the engine would not know unless we mark_dirty >>> b.backward() # RuntimeError: one of the variables needed for gradient >>> # computation has been modified by an inplace operation
- mark_non_differentiable(*args)[source]#
Mark outputs as non-differentiable.
This should be called at most once, in either the
setup_context()orforward()methods, and all arguments should be tensor outputs.This will mark outputs as not requiring gradients, increasing the efficiency of backward computation. You still need to accept a gradient for each output in
backward(), but it’s always going to be a zero tensor with the same shape as the shape of a corresponding output.- This is used e.g. for indices returned from a sort. See example::
>>> class Func(Function): >>> @staticmethod >>> def forward(ctx, x): >>> sorted, idx = x.sort() >>> ctx.mark_non_differentiable(idx) >>> ctx.save_for_backward(x, idx) >>> return sorted, idx >>> >>> @staticmethod >>> @once_differentiable >>> def backward(ctx, g1, g2): # still need to accept g2 >>> x, idx = ctx.saved_tensors >>> grad_input = torch.zeros_like(x) >>> grad_input.index_add_(0, idx, g1) >>> return grad_input
- save_for_backward(*tensors)[source]#
Save given tensors for a future call to
backward().save_for_backwardshould be called at most once, in either thesetup_context()orforward()methods, and only with tensors.All tensors intended to be used in the backward pass should be saved with
save_for_backward(as opposed to directly onctx) to prevent incorrect gradients and memory leaks, and enable the application of saved tensor hooks. Seetorch.autograd.graph.saved_tensors_hooks. See Extending torch.autograd for more details.Note that if intermediary tensors, tensors that are neither inputs nor outputs of
forward(), are saved for backward, your custom Function may not support double backward. Custom Functions that do not support double backward should decorate theirbackward()method with@once_differentiableso that performing double backward raises an error. If you’d like to support double backward, you can either recompute intermediaries based on the inputs during backward or return the intermediaries as the outputs of the custom Function. See the double backward tutorial for more details.In
backward(), saved tensors can be accessed through thesaved_tensorsattribute. Before returning them to the user, a check is made to ensure they weren’t used in any in-place operation that modified their content.Arguments can also be
None. This is a no-op.See Extending torch.autograd for more details on how to use this method.
Example:
>>> class Func(Function): >>> @staticmethod >>> def forward(ctx, x: torch.Tensor, y: torch.Tensor, z: int): >>> w = x * z >>> out = x * y + y * z + w * y >>> ctx.save_for_backward(x, y, w, out) >>> ctx.z = z # z is not a tensor >>> return out >>> >>> @staticmethod >>> @once_differentiable >>> def backward(ctx, grad_out): >>> x, y, w, out = ctx.saved_tensors >>> z = ctx.z >>> gx = grad_out * (y + y * z) >>> gy = grad_out * (x + z + w) >>> gz = None >>> return gx, gy, gz >>> >>> a = torch.tensor(1., requires_grad=True, dtype=torch.double) >>> b = torch.tensor(2., requires_grad=True, dtype=torch.double) >>> c = 4 >>> d = Func.apply(a, b, c)
- save_for_forward(*tensors)[source]#
Save given tensors for a future call to
jvp().save_for_forwardshould be called at most once, in either thesetup_context()orforward()methods, and all arguments should be tensors.In
jvp(), saved objects can be accessed through thesaved_tensorsattribute.Arguments can also be
None. This is a no-op.See Extending torch.autograd for more details on how to use this method.
Example:
>>> class Func(torch.autograd.Function): >>> @staticmethod >>> def forward(ctx, x: torch.Tensor, y: torch.Tensor, z: int): >>> ctx.save_for_backward(x, y) >>> ctx.save_for_forward(x, y) >>> ctx.z = z >>> return x * y * z >>> >>> @staticmethod >>> def jvp(ctx, x_t, y_t, _): >>> x, y = ctx.saved_tensors >>> z = ctx.z >>> return z * (y * x_t + x * y_t) >>> >>> @staticmethod >>> def vjp(ctx, grad_out): >>> x, y = ctx.saved_tensors >>> z = ctx.z >>> return z * grad_out * y, z * grad_out * x, None >>> >>> a = torch.tensor(1., requires_grad=True, dtype=torch.double) >>> t = torch.tensor(1., dtype=torch.double) >>> b = torch.tensor(2., requires_grad=True, dtype=torch.double) >>> c = 4 >>> >>> with fwAD.dual_level(): >>> a_dual = fwAD.make_dual(a, t) >>> d = Func.apply(a_dual, b, c)
- set_materialize_grads(value)[source]#
Set whether to materialize grad tensors. Default is
True.This should be called only from either the
setup_context()orforward()methods.If
True, undefined grad tensors will be expanded to tensors full of zeros prior to calling thebackward()andjvp()methods.Example:
>>> class SimpleFunc(Function): >>> @staticmethod >>> def forward(ctx, x): >>> return x.clone(), x.clone() >>> >>> @staticmethod >>> @once_differentiable >>> def backward(ctx, g1, g2): >>> return g1 + g2 # No check for None necessary >>> >>> # We modify SimpleFunc to handle non-materialized grad outputs >>> class Func(Function): >>> @staticmethod >>> def forward(ctx, x): >>> ctx.set_materialize_grads(False) >>> ctx.save_for_backward(x) >>> return x.clone(), x.clone() >>> >>> @staticmethod >>> @once_differentiable >>> def backward(ctx, g1, g2): >>> x, = ctx.saved_tensors >>> grad_input = torch.zeros_like(x) >>> if g1 is not None: # We must check for None now >>> grad_input += g1 >>> if g2 is not None: >>> grad_input += g2 >>> return grad_input >>> >>> a = torch.tensor(1., requires_grad=True) >>> b, _ = Func.apply(a) # induces g2 to be undefined
- set_output_grad_dtype(*dtypes)[source]#
Declare the gradient dtype for each of this Function’s outputs.
This should be called at most once, from either the
setup_context()orforward()methods. The number of declarations must match the number of returned values, and each argument corresponds positionally to the output at the same index.For each output, pass the dtype your backward should receive its gradient in:
Pass a
torch.dtypeand the engine guarantees the gradient handed to backward has that dtype. This is only valid for a differentiable Tensor output.Pass
Noneand the gradient is handed to backward with whatever dtype it already has. This is also the only valid choice for a non-Tensor or non-differentiable output, which has no gradient.Omit this call (or pass the output’s own dtype) and the gradient is handed to backward in the output’s dtype, which is the default.
For example:
>>> @staticmethod >>> def forward(ctx, x): >>> t1 = x.sin() >>> t2 = x.cos() >>> t3 = x.tan() >>> ctx.set_output_grad_dtype(torch.float32, t2.dtype, None, None) >>> return t1, t2, t3, "not a tensor"
This ensures that backward receives
t1’s gradient infloat32, keeps the default behavior fort2’s gradient viat2.dtype, passest3’s gradient through uncast withNone, and usesNoneas the placeholder for the trailing non-Tensor output.
- static setup_context(ctx, inputs, output)[source]#
There are two ways to define the forward pass of an autograd.Function.
Either:
Override forward with the signature
forward(ctx, *args, **kwargs).setup_contextis not overridden. Setting up the ctx for backward happens inside theforward.Override forward with the signature
forward(*args, **kwargs)and overridesetup_context. Setting up the ctx for backward happens insidesetup_context(as opposed to inside theforward)
See
torch.autograd.Function.forward()and Extending torch.autograd for more details.- Return type:
- static vjp(ctx, *grad_outputs)[source]#
Define a formula for differentiating the operation with backward mode automatic differentiation.
This function is to be overridden by all subclasses. (Defining this function is equivalent to defining the
vjpfunction.)It must accept a context
ctxas the first argument, followed by as many outputs as theforward()returned (None will be passed in for non tensor outputs of the forward function), and it should return as many tensors, as there were inputs toforward(). Each argument is the gradient w.r.t the given output, and each returned value should be the gradient w.r.t. the corresponding input. If an input is not a Tensor or is a Tensor not requiring grads, you can just pass None as a gradient for that input.The strides of the gradients passed to
backward()are undefined: they are not guaranteed to be contiguous or to match the strides of the corresponding forward outputs, so implementations must not assume a particular memory layout.The context can be used to retrieve tensors saved during the forward pass. It also has an attribute
ctx.needs_input_gradas a tuple of booleans representing whether each input needs gradient. E.g.,backward()will havectx.needs_input_grad[0] = Trueif the first input toforward()needs gradient computed w.r.t. the output.- Return type:
- static vmap(info, in_dims, *args)[source]#
Define the behavior for this autograd.Function underneath
torch.vmap().For a
torch.autograd.Function()to supporttorch.vmap(), you must either override this static method, or setgenerate_vmap_ruletoTrue(you may not do both).If you choose to override this staticmethod: it must accept
an
infoobject as the first argument.info.batch_sizespecifies the size of the dimension being vmapped over, whileinfo.randomnessis the randomness option passed totorch.vmap().an
in_dimstuple as the second argument. For each arg inargs,in_dimshas a correspondingOptional[int]. It isNoneif the arg is not a Tensor or if the arg is not being vmapped over, otherwise, it is an integer specifying what dimension of the Tensor is being vmapped over.*args, which is the same as the args toforward().
The return of the vmap staticmethod is a tuple of
(output, out_dims). Similar toin_dims,out_dimsshould be of the same structure asoutputand contain oneout_dimper output that specifies if the output has the vmapped dimension and what index it is in.Please see Extending torch.func with autograd.Function for more details.