You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Place and size each copy of a multicast with the node it reaches - #156
A transfer that multicasts a tensor to readers on different cores gave every copy the transfer's whole set of destination cores, and sized every copy with the transfer's own inter-core tiling. So the copy for a reader left on one core (a residual Add on a vector core, a softmax exp) was also reserved on the cores of the other readers, and was sized as one slice of the other readers' split instead of the whole tile it reads.
DecisionSpace narrows each copy's placements to the cores of the node it reaches; a copy for another transfer or an edge keeps the transfer's.
A chosen route only requires a copy on the targets where that copy can live.
Workload.get_tensor_single_core sizes a copy that feeds a computation node directly with that node's split; a copy staged for another transfer keeps the transfer's tiling. sliding_halo returns no halo for a copy its reader reads without a window.
Tests: in the attention block on the TPU-like quad core, each copy of a multicast can only be placed with its reader, and the copy for the unsplit exp holds and receives the whole tile. Both fail without the change.
To regenerate baseline:python scripts/analysis/render_metrics_comment.py --update-baseline
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A transfer that multicasts a tensor to readers on different cores gave every copy the transfer's whole set of destination cores, and sized every copy with the transfer's own inter-core tiling. So the copy for a reader left on one core (a residual Add on a vector core, a softmax exp) was also reserved on the cores of the other readers, and was sized as one slice of the other readers' split instead of the whole tile it reads.
DecisionSpacenarrows each copy's placements to the cores of the node it reaches; a copy for another transfer or an edge keeps the transfer's.Workload.get_tensor_single_coresizes a copy that feeds a computation node directly with that node's split; a copy staged for another transfer keeps the transfer's tiling.sliding_haloreturns no halo for a copy its reader reads without a window.Tests: in the attention block on the TPU-like quad core, each copy of a multicast can only be placed with its reader, and the copy for the unsplit exp holds and receives the whole tile. Both fail without the change.
Stacked on #155.