You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Give a node the fused loop of a dim it steps along through a strided coupling - #158
When a group is fused along the output rows of a stride-2 reader (a max pool), its producers reach that axis only through the coupling 2*o + r. The steady-state loops were added to a node only if the fused dim was one of its plain dims, so the conv and ReLU before the pool got no loop: they were costed for one of the iterations only and their outputs were held whole on chip.
_add_temporal_iteration_variables now matches a fused dim against the dim each node dim steps along fastest (Workload.leading_dim).
Test: a conv, ReLU and stride-2 max pool fused along the pool's rows loop with it on every node. It fails without the change. The stride-2 conv pair (S5) now costs conv1 on every iteration: 65,474 cycles on the Eyeriss-like quad core instead of 44,300, where conv1 was charged a quarter of its tile.
To regenerate baseline:python scripts/analysis/render_metrics_comment.py --update-baseline
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When a group is fused along the output rows of a stride-2 reader (a max pool), its producers reach that axis only through the coupling
2*o + r. The steady-state loops were added to a node only if the fused dim was one of its plain dims, so the conv and ReLU before the pool got no loop: they were costed for one of the iterations only and their outputs were held whole on chip._add_temporal_iteration_variablesnow matches a fused dim against the dim each node dim steps along fastest (Workload.leading_dim).Test: a conv, ReLU and stride-2 max pool fused along the pool's rows loop with it on every node. It fails without the change. The stride-2 conv pair (S5) now costs conv1 on every iteration: 65,474 cycles on the Eyeriss-like quad core instead of 44,300, where conv1 was charged a quarter of its tile.
Stacked on #157.