Skip to content

decouple IDX format from raw tensor #42

Description

@lgarithm

The IDX format only defined encodings for: u8, i8, i16, i32, f32, f64.
The new encoding format should support more scalar data types, and should be extendable.

Activity

  1. lgarithm commented on Nov 17, 2019

    @lgarithm
    CollaboratorAuthor

    The new encoding scheme should support all scalar data types given by the combination:

    (category, bytes, bits-per-byte, variant)

    where

    category := unsigned | signed | floating point
    bytes := 1 | 2 | 4 | 8 or other values if necessary
    bits-per-byte := 8 or other values if necessary
    variant := 0 | or other values if present

    the variant component is useful to distinguish IEEE 754 float16 and bfloat16.

    string type can be also encoded by the encoding scheme by adding a new category.

    type category bytes bits-per-byte variant
    u8 unsigned 1 8 0
    u16 unsigned 2 8 0
    u32 unsigned 4 8 0
    u64 unsigned 8 8 0
    u128 unsigned 16 8 0
    i8 signed 1 8 0
    i16 signed 2 8 0
    i32 signed 4 8 0
    i64 signed 8 8 0
    i128 signed 16 8 0
    f16 (IEEE 754) floating point 2 8 0
    f16 (bfloat16) floating point 2 8 1
    f32(IEEE 754) floating point 4 8 0
    f64(IEEE 754) floating point 8 8 0
    f128 ? floating point 16 8 ?
    basic_string<char> str sizeof(basic_string<char>) 8 ?

    Binary represent of the encoding:

    |01234567|01234567|
    |AAABBBBB|CCCCDDDD|
    
    component bits value range common values
    category 3 0 ~ 7 1,2,3,4, ...
    bytes 5 0 ~ 31 1,2,4,8,12,16
    bits-per-byte 4 0 ~ 15 8
    variant 4 0 ~ 15 0,1
  2. pinned this issue on Dec 16, 2019
  3. self-assigned this
    on Dec 16, 2019
  4. lgarithm commented on Dec 17, 2019

    @lgarithm
    CollaboratorAuthor
    template <typename R>
    constexpr uint32_t basic_scalar_category()
    {
        if (std::is_integral<R>::value) {
            if (std::is_unsigned<R>::value) { return 0; }
            if (std::is_signed<R>::value) { return 1; }
        }
        if (std::is_floating_point<R>::value) { return 2; }
        return -1;
    }
    
    // template <typename R, typename V = uint32_t>
    // struct basic_scalar_category;
    
    template <typename R, typename V = uint32_t>
    class basic_scalar_encoding;
    
    template <typename R, typename V>
    class basic_scalar_encoding
    {
        static constexpr V category = basic_scalar_category<R>();
        static constexpr V byte_num = sizeof(R);
        static constexpr V byte_size = CHAR_BIT;
    
      public:
        static constexpr V value = (category << 16) | (byte_num << 8) | byte_size;
    };
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions