- 18 Oct, 2023 2 commits
-
-
Yi-Yen Chung authored
Co-authored-by:
Décio Luiz Gazzoni Filho <decio@decpp.net> Co-authored-by:
Michael R. Crusoe <michael.crusoe@gmail.com>
-
Yi-Yen Chung authored
* [NEON] Add vabal_{s/u}{8/16/32} * [NEON] Add vabal_high_{s/u}{8/16/32} * [NEON] Add all vcale* intrinsics (9) * [NEON] Add all vcalt intrinsics (9) * [NEON] Add vcreate_f16 * [NEON] Add vreinterpret_u64_f16 * [NEON] Add vcvth_f16_s16 and vcvth_f16_u16 * [NEON] Add vduph_lane_f16, vdup_lane_f16, and vdupq_lane_f16 * [NEON] Add vext_f16 * [NEON] Add 16 vcvt{q}_n_* intrinsics * [Fix] Correct function input parameters * [NEON] Add 6 vcvtn_{s/u}{16/32/64}_f{*} intrinsics * [Fix] Correct vdup_lane_f16 and vdupq_lane_f16. * [Fix] Correct function input parameters. * [NEON] Add 24 vcvt{q}_n_* intrinsics * [NEON] Add all vcvtn* intrinsics * [NEON] Add vfmah_f16 and vfma_f16 * [NEON] Add vfma_n_f16 and vfmaq_n_f16 * [NEON] Add vmulh_f16 * [NEON] Add fma_lane related intrinsics. * [NEON] Add 5 vmul* related intrinsics vmulh_lane_f16, vmulh_laneq_f16, vmul_lane_f16, vmul_laneq_f16, vmulq_laneq_f16. * [NEON] Add neg related intrinsics. * [NEON] Add all fms, fms_n, and fms_lane intrinsics * [NEON] Add types float16x{4/8}x{2/3/4} * [NEON] Add 9 vld1 related intrinsics * [Fix] Modified wrong rounding implementation. Modified wrong implementation "Ties to Away" to "rounding to nearest with ties to Away" add.h: Remove redundant code. * [Fix] Fix wrong intrinsic alias names. * [Refactor] Remove redundant functions. * [NEON] Add 45 ld2 related intrinsics one ld2_f16, twenty-two ld2_lane series, and twenty-two ld2_dup series. * [NEON] Add ld3_dup, ld3_lane, and ld4_dup * [NEON] Add vld3_f16 and vld4_f16. * [NEON] Add vld{3/4}_{dup/lane} series intrinsics * [NEON] Add mla_{high}_lane series intrinsics * [NEON] Add qdmlal_{high}_{lane} series intrinsics. * [NEON] Add qdmlal_lane and qdmlal_n series intrinsics * [NEON] Add mls_lane and mlsl_high_lane series intrinsics * [NEON] Add 22 qdmlsl series intrinsics * [NEON] Add 10 qdmull_* series intrinsics * [NEON] Add 3 qdmulh series intrinsics * [Fix] Fix wrong function name. * [Fix] Correct the wrong alias function name. * [NEON] Add qdmullh_lane{q}_s{16/32} related intrinsics * [NEON] Add qdmull_n and qdmull_high_lane series intrinsics * [Fix] Add conditions for fp16 intrinsics * [Hack] Skip functions that trigger compiler bugs.
-
- 17 Oct, 2023 1 commit
-
-
Chi-Wei Chu authored
* [Neon] Add vcadd_rot270_f{16/32} and vcaddq_rot270_f{16/32/64} * [Neon] Add vcadd_rot90_f{16/32} and vcaddq_rot90_f{16/32/64} * [Neon] Add vcmla_lane_f{16/32} and vcmla_laneq_f{16/32} and vcmlaq_lane_f{16/32} and vcmlaq_laneq_f{16/32} * [Neon] Add vcmla_rot90_lane_f{16/32} and vcmla_rot90_laneq_f{16/32} and vcmlaq_rot90_lane_f{16/32} and vcmlaq_rot90_laneq_f{16/32} * [Neon] Add vcmla_rot180_lane_f{16/32} and vcmla_rot180_laneq_f{16/32} and vcmlaq_rot180_lane_f{16/32} and vcmlaq_rot180_laneq_f{16/32} * [Neon] Add vcadd_rot270_f{16/32} and vcaddq_rot270_f{16/32/64}
-
- 16 Oct, 2023 2 commits
-
-
Michael R. Crusoe authored
Adapted from https://github.com/DLTcollab/sse2neon/pull/6 and https://github.com/DLTcollab/sse2neon/pull/559
-
Michael R. Crusoe authored
-
- 13 Oct, 2023 3 commits
-
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
Inspired by https://github.com/DLTcollab/sse2neon/pull/579
-
George Vinokhodov authored
-
- 06 Oct, 2023 1 commit
-
-
Yi-Yen Chung authored
* [NEON] Add vmulq_f16 and vmul_f16. * [NEON] Add vmulq_f16 and vmul_f16 test. * [NEON] Add vget_lane_f16 and vgetq_lane_f16. * [NEON] Add vsubh_f16 and vsub_f16. * [NEON] Add vextq_f16. * [NEON] Add vget_low_f16. * [NEON] Add vmulq_lane_f16. * [NEON] Add vmul_n_f16. * [NEON] Add vget_high_f16. * [NEON] Add vsetq_lane_f16. * [NEON] Add vcombine_f16. * [NEON] Add vcvtaq_s32_f32, vcvtas_s32_f32, and vcvta_s32_f32. * [NEON] Add vpadd_f16. * [NEON] Add vuzp1_f16. * [NEON] Add vuzp2_f16. * [NEON] Add vmaxq_f16 and vmax_f16. * [NEON] Add vcvtas_u32_f32, vcvta_u32_f32, vcvtaq_u32_f32. * [NEON] Add type simde_float16x8x2_t. * [NEON] Add vld2q_f16. * [NEON] Add vld1q_dup_f16. * [NEON] Add vpmax_f16. * [NEON] Add vrsqrtsq_f16, vrsqrtsh_f16, vrsqrts_f16. * [NEON] Add vcgtq_f16, vcgt_f16, vcgth_f16. * [NEON] Add vdiv_f32 and vdivq_f32. * [NEON] Add vrecps_f16 and vrecpsq_f16. * [NEON] Add vset_lane_f16. * [NEON] Add vrecpe_f16, vrecpeq_f16. * [NEON] Add vfmaq_f16. * [NEON] Add vabsq_f16 and vabs_f16. * [NEON] Add vcltq_f16 and vclth_f16. * [NEON] Add vmin_f16 and vminq_f16. * [NEON] Add vclt_f16. * [NEON] Add vclezq_f16, vclez_f16, and vclezh_f16. * [NEON] Add vzip_f16 and vzipq_f16. * [NEON] Add vzip1_f16 and vzip1q_f16. * [NEON] Add vzip2_f16 and vzip2q_f16. * [NEON] Add vst2_f16 and vst2q_f16. * [NEON] Add type simde_float16x4x2. * [NEON] Add 16 intrinsics of vreinterpret series. * [SIMDE] Add sqrtl() in simde_math_sqrtl. * [NEON] Add vrndnq_f16, vrndns_f16, and vrndn_f16. * [NEON] Add sqrt in meson.build. * [NEON] Add 7 sqrt related intrinsics. * [Fix] Add the judge whether define sqrt() or not. * [NEON] Add vrsqrteq_f16, vrsqrte_f16, and vrsqrteh_f16. * [NEON] Add vqrshrnh_n_s16 and vqrshrnh_n_u16. * [NEON] Add vqrshrunh_n_s16. * [Fix] Add new conditions for fp16 intrinsics. * [License] Add Copyright.
-
- 02 Oct, 2023 1 commit
-
-
Michael R. Crusoe authored
Closes: #926 Co-authored-by:clin99 <34017491+clin99@users.noreply.github.com>
-
- 28 Sep, 2023 2 commits
- 26 Sep, 2023 2 commits
- 18 Aug, 2023 1 commit
-
-
Chris Bielow authored
when using `/arch=AVX` or similar then SSE4.2 does not get activated (on MSVC at least, but should be similar on g++/clang) This PR fixes this.
-
- 16 Jun, 2023 1 commit
-
-
Michael R. Crusoe authored
Fixes errors like FAILED: test/x86/avx512/cast-emul-c (also load, loadu, max, min, permutexvar, reduce, set, setzero, set1, storeu, f16c) clang-17 -o test/x86/avx512/cast-emul-c test/x86/avx512/cast-emul-c.p/cast.c.o -Wl,--as-needed -Wl,--no-undefined --target=riscv64-linux-gnu -Wl,--start-group -lm -Wl,--end-group /usr/bin/riscv64-linux-gnu-ld: test/x86/avx512/cast-emul-c.p/cast.c.o: in function `simde_test_x86_assert_equal_f16x32_': /opt/simde/riscv64-clang-17/../../../usr/local/src/simde/test/x86/avx512/test-avx512.h:14: undefined reference to `__truncsfhf2' FAILED: test/x86/avx512/cmp-emul-c (also reduce) clang-17 -o test/x86/avx512/cmp-emul-c test/x86/avx512/cmp-emul-c.p/cmp.c.o -Wl,--as-needed -Wl,--no-undefined --target=riscv64-linux-gnu -Wl,--start-group -lm -Wl,--end-group /usr/bin/riscv64-linux-gnu-ld: test/x86/avx512/cmp-emul-c.p/cmp.c.o: in function `simde_mm512_cmp_ph_mask': /opt/simde/riscv64-clang-17/../../../usr/local/src/simde/simde/x86/avx512/cmp.h:741: undefined reference to `__extendhfsf2' /usr/bin/riscv64-linux-gnu-ld: /opt/simde/riscv64-clang-17/../../../usr/local/src/simde/simde/x86/avx512/cmp.h:741: undefined reference to `__extendhfsf2' /usr/bin/riscv64-linux-gnu-ld: /opt/simde/riscv64-clang-17/../../../usr/local/src/simde/simde/x86/avx512/cmp.h:741: undefined reference to `__extendhfsf2' /usr/bin/riscv64-linux-gnu-ld: /opt/simde/riscv64-clang-17/../../../usr/local/src/simde/simde/x86/avx512/cmp.h:741: undefined reference to `__extendhfsf2' /usr/bin/riscv64-linux-gnu-ld: /opt/simde/riscv64-clang-17/../../../usr/local/src/simde/simde/x86/avx512/cmp.h:741: undefined reference to `__extendhfsf2' /usr/bin/riscv64-linux-gnu-ld: test/x86/avx512/cmp-emul-c.p/cmp.c.o:/opt/simde/riscv64-clang-17/../../../usr/local/src/simde/simde/x86/avx512/cmp.h:741: more undefined references to `__extendhfsf2' follow clang: error: linker command failed with exit code 1 (use -v to see invocation)
-
- 15 Jun, 2023 1 commit
-
-
Michael R. Crusoe authored
-
- 13 Jun, 2023 3 commits
-
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
sse/sse2: added ".f16" float16 accessor for m128 & m128i mm{256,512}_cvt{ph_ps,ps_ph}: use .f16 instead of converting from u16 mm512_cvt{epi8_epi16,epi32_ps,cvtph_ps}: added _mm256_cvt* native fallbacks
-
- 06 Jun, 2023 1 commit
-
-
Michael R. Crusoe authored
-
- 01 Jun, 2023 4 commits
-
- 30 May, 2023 1 commit
-
-
Simon Gene Gottlieb authored
-
- 28 May, 2023 4 commits
-
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
- 22 May, 2023 1 commit
-
-
easyaspi314 authored
Clang incorrectly sets the alignment bit on 32-bit for the multiple load instructions. This results in a bus error on hardware.
-
- 21 May, 2023 7 commits
-
-
easyaspi314 authored
Similar to _mm_shuffle_epi8, qtbl/qtbx now split up to multiple shuffles on A32V7. With multi-vector shuffles, two tables are created, the indexes are subtracted, and either TBX'd on the result of the other table or ORR'd together (TBX is mandatory for qtbx, but ORR is chosen when the lookups would not be theoretically dual issued).
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
Caught by clang-14+
-
Michael R. Crusoe authored
-
- 20 May, 2023 2 commits
-
-
Michael R. Crusoe authored
-
Michael R. Crusoe authored
-