15 Jul, 2021 1 commit
- default iterator hack for blockwise copy · c94efafd
  Chao Liu authored 4 years ago
  
  c94efafd
09 Jul, 2021 1 commit
- fix bug: config for ThreadwiseDynamicTensorSliceTransfer_v2 (#46) · 1c1b56fe
  Chao Liu authored 4 years ago
  
  Unverified
  
  1c1b56fe
08 Jul, 2021 4 commits
- Create README.md (#45) · 4682d070
  Chao Liu authored 4 years ago
```
* Create README.md
```
  Unverified
  
  4682d070
- Tweak (#44) · aafb5eb1
  Chao Liu authored 4 years ago
```
* tweak
```
  Unverified
  
  aafb5eb1
- Update default launch bounds (#43) · 2f82cfb1
  Chao Liu authored 4 years ago
```
* update default launch bounds
```
  Unverified
  
  2f82cfb1
- Deprecate static kernel (#42) · 81c942cd
  Chao Liu authored 4 years ago
```
* deprecate static kernels
```
  Unverified
  
  81c942cd
05 Jul, 2021 1 commit

DL GEMM fp32/fp16/int8 (#41) · b8b2d0a6

Chao Liu authored 4 years ago

* add threadwise copy the copy a tensor in one copy, added kpack to DL GEMM

* add kpack into fwd v4r5 nchw fp32

b8b2d0a6

01 Jul, 2021 2 commits

fix complain about divide by zero (#40) · 11ec07e9
Chao Liu authored 4 years ago

Unverified

11ec07e9

xdlops_v4r4_fwd fp32/fp16 (#34) · 3835318c

zjing14 authored 4 years ago


* create files for xdlops

* working on blockwise_gemm_xdlops

* add KReduction

* add m/n repeats

* add 2x2 pipeline

* added 128x128 wavegemm

* use StaticBuffer of vector_type

* break vector type to blk_size

* add kpack into xldops_gemm and blockwise_gemm

* abroadcast only

* add fp32 mfma instructions

* adding fp16 mfma

* pack half4_t

* rename kperwave to kpack

* add 32x32x8fp16

* add fp16 mfma

* clean code

* clean code

* V4r4 xdlops kpack (#35)

* add kpack with incorrect results

* bug fix for make_dynamic_naive_tensor_descriptor_aligned_v2

* add 1x1 kernel

* add gridwise_gemm_v2 - single_buffer

* enabled dwordx4 for fp16
Co-authored-by: Chao Liu <chao.liu2@amd.com>

* refactor fwd-v4r4-xdlops

* add v4r4-nhwc-xdlop

* improve some perf of nhwc and nchw by tuning parameters, and change scheuduling in gridwise-gemm loop

* tweak scheduling in gridwise gemm

* add v4r3 with a single output copy

* init commit: output with slice win

* adding sliceWin

* add multiple repeats pattern

* starting adding bwd-v4r1-xdlops

* use tuple as SrcBuffer

* adding bwd-data v4r1 nhwc xdlops

* fix bug in make_dynamic_naive_tensor_descriptor_aligned_v2()

* fix bug in host bwd-data conv

* initial implementation of bwd-data v4r1 nhwc xdlops

* add launch bound flags

* enable launch bound

* add m/nrepeat=4

* tweak bwd-data v4r1 nhwc xdlops

* added bwd-data v4r1 nhwc xlops with output A and weight B

* add fwd-v4r4 nhwc xdlops, A input, B weight, C output
Co-authored-by: Chao Liu <chao.liu2@amd.com>

3835318c

24 Jun, 2021 1 commit

Add online compilation for dynamic kernels (#37) · 1685048a

Qianfeng authored 4 years ago

* Add online-compiling facility

* Synchronize from fwd-v4r5 and implement host interfaces to call conv-fwd v4r4/v4r5 using on-line compiling method

* Tiny adjustment to time reporting

* Use object assignment to replace explicit bytes copying in the first kernel of v4r4/v4r5

* Use single thread to assign descriptor object to device memory

* Adjust to the workload assignment of the two kernels of v4r4 (experimental)

* Revert "Adjust to the workload assignment of the two kernels of v4r4 (experimental)"

This reverts commit eb384614.

* Update to make constexpr for generating descriptor types in kernel 2 of dynamic conv-fwd v4r4

* Update to dynamic conv-fwd v4r4 online-compiling

* Update to dynamic conv-fwd v4r5 online-compiling (result not accurate)

* Tiny update to driver/CMakeLists.txt

* clang-format

* Tiny comments change

* Add env OLC_DUMP_SAVE_TMP_DIR to support saving of temperary dir

* Fwd v4r5 olc perf (#39)

* added hip-clang flags that fix perf issue of online compilation

* fix bug for olc fwd-v4r5-nchw

* Move constexpr and type reference statements out of the function body in conv-fwd v4r4/v4r5 kernel wrapper

* Remove printing in hip_build_utils.cpp

* Update to root CMakeLists.txt

* Revert "Move constexpr and type reference statements out of the function body in conv-fwd v4r4/v4r5 kernel wrapper"

This reverts commit 3d2c5d8e

.
Co-authored-by: Chao Liu <chao.liu2@amd.com>
Co-authored-by: Chao Liu <lc.roy86@gmail.com>
Co-authored-by: root <root@dc-smc-18.amd.com>

1685048a

19 Jun, 2021 1 commit
- pass-by-void-pointer for gridwise_dynamic_gemm_v1r2 (#38) · d2315b0d
  Chao Liu authored 4 years ago
```
* pass-by-void-pointer for gridwise_dynamic_gemm_v1r2

* use pass-by-value by default
```
  Unverified
  
  d2315b0d
10 Jun, 2021 1 commit

Restructure gridwise and blockwise GEMM, add tensor contraction and FWD-v4r5 (#36) · 30072aec

Chao Liu authored 4 years ago

* experimenting magic number division

* overhauling fwd-v4r4 to clearly reflect transformation graph

* added fwd-v4r5

* bug fix for make_dynamic_naive_tensor_descriptor_aligned_v2

* bug fix and added sanity-check in transform_dynamic_tensor_descriptor

* added conv_driver_v2

30072aec

12 May, 2021 2 commits
- reorganize some files (#33) · 71d6b19d
  Chao Liu authored 4 years ago
  
  Unverified
  
  71d6b19d
- Use DynamicBuffer instead of raw pointer (#32) · 78b987fb
  Chao Liu authored 4 years ago
```
* Use DynamicBuffer to hold raw pointer (to global and LDS memory)

* add workaround for compiler issue (inefficient ISA) of ds_write for int8x4, int8x8, int8x16
```
  Unverified
  
  78b987fb
11 May, 2021 1 commit

No raw index calculation (#31) · 01055d95

Chao Liu authored 4 years ago


* Replace most raw index calculation to coordinate transformation
* Overhaul blockwise and threadwise GEMM
* Overhaul driver for gridwies GEMM kernel
Co-authored-by: Jing Zhang <jizhan@amd.com>

01055d95

28 Apr, 2021 1 commit
- Use Tuple and vector_type instead of Array for holding tensor data (#30) · d075adf1
  Chao Liu authored 4 years ago
```
* replacing array with tuple and vector for tensor data
```
  Unverified
  
  d075adf1
13 Apr, 2021 2 commits
- Overhaul vector_type and use real vector for int8x4_t instead of aliasing from int32_t (#29) · e4790c25
  Chao Liu authored 4 years ago
```
* overhaul vector_type, make int8x4_t real vector instead of aliasing from int32_t
```
  Unverified
  
  e4790c25
- Initial implementation of magic number division and "Merge" transformation that use it (#28) · 3bf52e60
  Chao Liu authored 4 years ago
```
* initial implementation for magic number division and DynamicMerge_v2_magic_division that uses it

* turn off DynamicMerge_v2_magic_division that use magic number division by default
```
  Unverified
  
  3bf52e60
07 Apr, 2021 1 commit
- Hybrid direct + implicit GEMM forward convolution NCHWc v5r1 (#25) · 792a20fa
  zjing14 authored 4 years ago
```
* Hybrid direct + implicit GEMM forward convolution NCHWc v5r1. Input tensor bypass LDS. Support fp32/fp16/int8
```
  Unverified
  
  792a20fa
06 Apr, 2021 2 commits
- Fix performance issue when passing tensor descriptor from host to kernel by void pointers (#27) · d2217f30
  Chao Liu authored 4 years ago
```
* use address_space(4) in kernel signature to fix performance issue when passing tensor descriptor from host to kernel by (void) pointers

* remove passing by pointer* option (only use pass by value or void*)
```
  Unverified
  
  d2217f30
- bug fix for buffer resource setting (#26) · 6a5ea493
  zjing14 authored 4 years ago
  
  Unverified
  
  6a5ea493
25 Mar, 2021 1 commit

Dynamic tensor descriptor (#24) · fcbb9788

Chao Liu authored 4 years ago


* support dynamic tensor descriptor

* use buffer load OOB feature for padding case

* add navi support

* add int8x4 inference kernel
Co-authored-by: Chao Liu <chao@ixt-rack-81.local.lan>
Co-authored-by: Jing Zhang <jizhan@amd.com>

fcbb9788

06 Aug, 2020 1 commit

Bwd Data NHWC (#22) · bbcb67d0

Chao Liu authored 4 years ago

* fix buffer_store bug
* remove obsolete kernels
* add bwd-data-v5r1-nhwc

bbcb67d0

29 Jul, 2020 1 commit

Improve buffer address for out of bound check (#21) · ac62d13e

Chao Liu authored 4 years ago

* Use buffer load built-in OOB check. buffer size is limited to 2GB.
* buffer APIs use combined wave and thread offset
* use uint32_t for addr shift in buffer addressing

ac62d13e

24 Jun, 2020 1 commit

Code clean up (#20) · 5c7cec11

Chao Liu authored 5 years ago


* tuning para,

* testing on v100

* add fp16

* remove deprecated tensor descriptor

* sync with miopen

* update build script
Co-authored-by: Jing Zhang <jizhan@amd.com>

5c7cec11

18 Feb, 2020 1 commit
- MIOpen integration (#15) · 7d09790a
  Chao Liu authored 5 years ago
```
* renaming
```
  Unverified
  
  7d09790a
17 Feb, 2020 1 commit
- MIopen integration (#13) · 1a66e35b
  Chao Liu authored 5 years ago
```
* update for miopen integration: cosmetic refactor
```
  Unverified
  
  1a66e35b
27 Jan, 2020 1 commit
- Update for recent MIOpen integration (#11) · 3406a114
  Chao Liu authored 5 years ago
```
* update for MIOpen integration
```
  Unverified
  
  3406a114
20 Jan, 2020 1 commit

Added bwd data v3r1 v4r1, tweaking v1 (#10) · c5da0377

Chao Liu authored 5 years ago

* Added bwd data v3r1: breaking down compute into a series of load balanced GEMM, and launch in a single kernel
* Added bwd data v4r1: like v3r1, but launch GEMMs in multiple kernels
* Tweaked v1r1  and v1r2 (atomic) on AMD GPU

c5da0377

05 Dec, 2019 1 commit
- update implicit GEMM forward v4r4 to use gridwise gemm (#9) · e2b4c5b4
  Chao Liu authored 5 years ago
```
* updated fwd v4r4 to use gridwise gemm
* updated gridwise gemm api calls in bwd-data v1r1 and v2r1
```
  Unverified
  
  e2b4c5b4
03 Dec, 2019 2 commits
- fixed faulty padding API calls (#8) · 19a93dac
  Chao Liu authored 5 years ago
  
  Unverified
  
  19a93dac
- backward data (#7) · 8f5f6496
  Chao Liu authored 5 years ago
```
* enabled atomic add in tensor copy
* added gridwise GEMM
* added backward data conv using GEMM + atomic
* added backward data conv using GEMM, no atomic
```
  Unverified
  
  8f5f6496
04 Nov, 2019 2 commits
- remove dead file (#6) · 31ded4ac
  Chao Liu authored 5 years ago
  
  Unverified
  
  31ded4ac
- MIOpen integration: recent bug fixes from MIOpen (#5) · 562e1e27
  Chao Liu authored 5 years ago
  
  Unverified
  
  562e1e27
11 Oct, 2019 1 commit
- Refactor for MIOpen integration (#4) · 52c3fe05
  Chao Liu authored 5 years ago
```
Refactor, so can bring multi-index transformation and padding support into MIOpen
```
  Unverified
  
  52c3fe05
30 Sep, 2019 2 commits
- Merge pull request #3 from asroy/clean_up · 9aaeacc8
  Chao Liu authored 5 years ago
```
enable type conversion in ThreadwiseGenericTensorSliceCopy_v2r1 and BlockwiseGenericTensorSliceCopy_v2
```
  Unverified
  
  9aaeacc8
- enable type conversion in blockwise copy v2 and threadwise copy v2r1 · cf218184
  Chao Liu authored 5 years ago
  
  cf218184
27 Sep, 2019 3 commits
- tweaking · 012d3a07
  Chao Liu authored 5 years ago
  
  012d3a07
- tweaking · 14315b72
  Chao Liu authored 5 years ago
  
  14315b72
- debugging · ebe38f3d
  Chao Liu authored 5 years ago
  
  ebe38f3d

Menu