Blog Archive

Friday, May 22, 2020

Some tricks to speed up your C++ code writing speed


3 Simple C++17 Features That Will Make Your Code Simpler

Published June 19, 2018
This article is a guest post written by guest author jft.
C++17 has brought a lot of features to the C++ language. Let’s dig into three of them that help make coding easier, more concise, intuitive and correct.
We’ll begin with Structured Bindings. These were introduced as a means to allow a single definition to define multiple variables with different types. Structured bindings apply to many situations, and we’ll see several cases where they can make code more concise and simpler.
Then we’ll see Template Argument Deduction, which allows us to remove template arguments that we’re used to typing, but that we really shouldn’t need to.
And we’ll finish with Selection Initialization, which gives us more control about object scoping and lets us define values where they belong.
So let’s start with structured bindings.

Structured Bindings

Structured Bindings allow us to define several objects in one go, in a more natural way than in the previous versions of C++.

From C++11 to C++17

This concept is not new in itself. Previously, it was always possible to return multiple values from a function and access them using std::tie.
Consider the function:
This returns three variables all of different types. To access these from a calling function prior to C++17, we would need something like:
Where the variables have to be defined before use and the types known in advance.
But using Structured Bindings, we can simply do this as:
which is a much nicer syntax and is also consistent with modern C++ style using auto almost whenever possible.
So what can be used with a Structured Binding initialization? Basically anything that is a compound type – struct, pair and tuple. Let’s see several cases where it can be useful.

Returning compound objects

This is the easy way to assign the individual parts of a compound type (such as a struct, pair etc) to different variables all in one go – and have the correct types automatically assigned. So let’s have a look at an example. If we insert into a map, then the result is a std::pair:
And if anyone is wondering why the types are not explicitly stated for pair, then the answer is Template Argument Deduction in C++17 – keep reading!
So to determine if the insert was successful or not, we could extract the info from what the insert method returned:
The problem with this code is that a reader needs to look up what .second is supposed to mean, if only mentally. But using Structured Bindings, this becomes:
Where itelem is the iterator to the element and success is of type bool, with true for insertion success. The types of the variables are automatically deduced from the assignment – which is much more meaningful when reading code.
As a sneak peek into the last section, as C++17 now has Selection Initialization, then we could (and probably would) write this as:
But more on this in a moment.

Iterating over a compound collection

Structured Bindings also work with range-for as well. So considering the previous mymap definition, prior to C++17 we would iterate it with code looking like this:
Or maybe, to be more explicit:
But Structured Bindings allow us to write it more directly:
The usage of the variables key and value are more instructive than entry.first and entry.second – and without requiring the extra variable definitions.

Direct initialization

But as Structured Bindings can initialize from a tuple, pair etc, can we do direct initialization this way?
Yes we can. Consider:
which defines variables a as type char with initial value ‘a’, i as type int with initial value 123 and b as type bool with initial value true.
Using Structured Bindings, this can be written as:
This will define the variables a, i, b the same as if the separate defines above had been used.
Is this really an improvement over the previous definition? OK, we’ve done in one line what would have taken three but why would we want to do this?
Consider the following code:
Both iss and name are only used within the for block, yet iss has to be declared outside of the for statement and within its own block so that the scope is limited to that required.
This is weird, because iss belongs to the for loop.
Initialization of multiple variables of the same type has always been possible. For example:
But what we’d like to write – but can’t – is:
With Structured Bindings we can write:
and
Which allows the variables iss and name (and i and ch) to be defined within the scope of the for statement as needed and also their type to be automatically determined.
And likewise with the if and switch statements, which now take optional Selection Initialization in C++17 (see below). For example:
Note that we can’t do everything with structured bindings, and trying to fit them in to every situation can make the code more convoluted. Consider the following example:
Here variable box is defined as type unsigned long and has an initial value returned from stoul(p). stoul(), for those not familiar with it, is a <string> function which takes a type std::string as its first argument (there are other optional ones – including base) and parses its content as an integral number of the specified base (defaults to 10), which is returned as an unsigned long value.
The type of variable bit is that of an iterator for boxes and has an initial value of .begin() – which is just to determine its type for auto. The actual value of variable bit is set in the condition test part of the if statement. This highlights a constraint with using Structured Bindings in this way. What we really want to write is:
But we can’t because a variable declared within an auto type specifier cannot appear within its own initializer! Which is kind of understandable.
So to sum up, the advantages of using Structured Bindings are:
  • a single declaration that declares one or more local variables
  • that can have different types
  • whose types are always deduced using a single auto
  • assigned from a composite type.
The drawback, of course, is that an intermediary (eg std::pair) is used. This needn’t necessarily impact upon performance (it is only done once at the start of the loop anyhow) as move semantics would be used where possible – but note that where a type used is non-moveable (eg like std::array) then this could incur a performance ‘hit’ depending upon what the copy operation involved.
But don’t pre-judge the compiler and pre-optimize code! If the performance isn’t as required, then use a profiler to find the bottleneck(s) – otherwise you are wasting development time. Just write the simplest / cleanest code that you can.

Template Argument Deduction

Put simply, Template Argument Deduction is the ability of templated classes to determine the type of the passed arguments for constructors without explicitly stating the type.
Before C++17, to construct an instance of a templated class we had to explicitly state the types of the argument (or use one of the make_xyz support functions).
Consider:
Here, p is an instance of the class pair and is initialized with values of 2 and 4.5. Or the other method of achieving this would be:
Both methods have their drawbacks. Creating “make functions” like std::make_pair is confusing, artificial and inconsistent with how non-template classes are constructed. std::make_pair, std::make_tuple etc are available in the standard library, but for user-defined types it is worse: you have to write your own make_… functions. Doh!
Specifying template arguments, as in:
should be unnecessary since they can be inferred from the type of the arguments – as is usual with template functions.
In C++17, this requirement for specifying the types for a templated class constructor has been abolished. This means that we can now write:
or
which is the logical way you would expect to be able to define p!
So considering the earlier function mytuple(). Using Template Argument Deduction (and auto for function return type), consider:
This is a much cleaner way of coding – and in this case we could even wrap it as:
There is more to it than that, and to dig deeper into that feature you can check out Simon Brand’s presentation about Template Argument Deduction.

Selection Initialization

Selection Initialization allows for optional variable initialization within if and switch statements – similar to that used within for statements. Consider:
Here the scope of a is limited to the for statement. But consider:
Here variable a is used only within the if statement but has to be defined outside within its own block if we want to limit its scope. But in C++17 this can be written as:
Which follows the same initialization syntax as the for statement – with the initialization part separated from the selection part by a semicolon (;). This same initialization syntax can similarly be used with the switch statement. Consider:
Which all nicely helps C++ to be more concise, intuitive and correct! How many of us have written code such as:
Where a before the second if hasn’t been initialized properly before the test (an error) but isn’t picked up by the compiler because of the earlier definition – which is still in scope as it isn’t defined within its own block. If this had been coded in C++17 as:
Then this would have been picked up by the compiler and reported as an error. A compiler error costs much less to fix than an unknown run-time problem!

C++17 helps making code simpler

In summary, we’ve seen how Structured Bindings allow for a single declaration that declares one or more local variables that can have different types, and whose types are always deduced using a single auto. They can be assigned from a composite type.
Template Argument Deduction allows us to avoid writing redundant template parameters and helper functions to deduce them. And Selection Initialization make the initialization in if and switch statements consistent with the one in for statements – and avoids the pitfall of variable scoping being too large.

References


Reference:
https://www.fluentcpp.com/2018/06/19/3-simple-c17-features-that-will-make-your-code-simpler/

Monday, March 16, 2020

How install ffmpeg plugin to make audacity support m4a


Step1: 
RECOMMENDED Installer Package for Windows: Lame_v3.99.3_for_Windows.exe 


Step2: 
FFmpeg RECOMMENDED ZIP OPTION: ffmpeg-win-2.2.2.zip


Step 3: 
Edit->preference->Library-> locate the path for ffmpeg:
for example:
OneDrive\tools\ffmpeg-win-2.2.2\avformat-55.dll


Reference:
[1] https://lame.buanzo.org/#lamewindl
[2] https://videoconverter.wondershare.com/convert-mp3/convert-m4a-to-mp3-audacity.html

Tuesday, March 3, 2020

Bitwise operations cheat sheet


Recommendations and additions to this cheat sheet are welcome.
This cheat sheet is mostly suitable for most common programming languages, but the target usage is C/C++ on x86 platform.
Bitmap i is unsigned 32 bit integers. For 64 bit operands, the suffix L should be added to integer literals, e.g. 1 should be 1L.
All calculation demonstration is done with 8 bits integers for readability.

Truth table

AND0 0 | 0
0 1 | 0
1 0 | 0
1 1 | 1OR0 0 | 0
0 1 | 1
1 0 | 1
1 1 | 1XOR0 0 | 0
0 1 | 1
1 0 | 1
1 1 | 0

Operators

AND         &
OR          |
NOT         ~
XOR         ^
Left shift  <<
Right shift >>

Get a bit

(i >> n) & 1

Set a bit to 1

i | (1 << n)

Set a bit to 0

i & ~(1 << n)

Store a bit

The bit to be stored is v which is either 0 or 1.
(i & ~(1 << n)) | (v << n)

Toggle a bit

i ^ (1 << n)

Get least significant bit

i & -i
Note: this gives you really the lowest bit but not the index of the lowest bit.
-i is equivalent to ~i + 1
~i    10100111
~i+1  10101000
i     01011000
i&-i  00001000

Get most significant bit

unsigned int get_msb(unsigned int i){
  i |= i >> 1;
  i |= i >> 2;
  i |= i >> 4;
  i |= i >> 8;
  i |= i >> 16;
  return (i + 1) >> 1;
}
How it works:
i         01000010
i|=i>>1   01100011
i|=i>>2   01111011
i|=i>>4   01111111
...
i|=i>>16  01111111
i+1       10000000
(i+1)>>1  01000000

Get index of most significant bit

inline unsigned int get_bit_index(const unsigned int i){
  unsigned int r;
  asm ( "bsr %1, %0\n"
      : "=r"(r)
      : "r" (i)
  );
  return r;
}
Yah, just one instruction. This instruction supports 16/32/64 bit integers, source and destination type must have the same size.

Change endianess

Convert from big-endian to little-endian or vice-versa.
Well you shouldn’t need to handcraft this function but anyway FYR:
((i>>24) & 0xFF)    |  // Move byte 3 to byte 0
((i<<8) & 0xFF0000) |  // Move byte 1 to byte 2
((i>>8) & 0xFF00)   |  // Move byte 2 to byte 1
((i<<24) & 0xFF000000) // Move byte 0 to byte 3

Bit reversal

static const unsigned char BitReverseTable256[] = 
{
  0x00, 0x80, 0x40, 0xC0, 0x20, 0xA0, 0x60, 0xE0, 0x10, 0x90, 0x50, 0xD0, 0x30, 0xB0, 0x70, 0xF0, 
  0x08, 0x88, 0x48, 0xC8, 0x28, 0xA8, 0x68, 0xE8, 0x18, 0x98, 0x58, 0xD8, 0x38, 0xB8, 0x78, 0xF8, 
  0x04, 0x84, 0x44, 0xC4, 0x24, 0xA4, 0x64, 0xE4, 0x14, 0x94, 0x54, 0xD4, 0x34, 0xB4, 0x74, 0xF4, 
  0x0C, 0x8C, 0x4C, 0xCC, 0x2C, 0xAC, 0x6C, 0xEC, 0x1C, 0x9C, 0x5C, 0xDC, 0x3C, 0xBC, 0x7C, 0xFC, 
  0x02, 0x82, 0x42, 0xC2, 0x22, 0xA2, 0x62, 0xE2, 0x12, 0x92, 0x52, 0xD2, 0x32, 0xB2, 0x72, 0xF2, 
  0x0A, 0x8A, 0x4A, 0xCA, 0x2A, 0xAA, 0x6A, 0xEA, 0x1A, 0x9A, 0x5A, 0xDA, 0x3A, 0xBA, 0x7A, 0xFA,
  0x06, 0x86, 0x46, 0xC6, 0x26, 0xA6, 0x66, 0xE6, 0x16, 0x96, 0x56, 0xD6, 0x36, 0xB6, 0x76, 0xF6, 
  0x0E, 0x8E, 0x4E, 0xCE, 0x2E, 0xAE, 0x6E, 0xEE, 0x1E, 0x9E, 0x5E, 0xDE, 0x3E, 0xBE, 0x7E, 0xFE,
  0x01, 0x81, 0x41, 0xC1, 0x21, 0xA1, 0x61, 0xE1, 0x11, 0x91, 0x51, 0xD1, 0x31, 0xB1, 0x71, 0xF1,
  0x09, 0x89, 0x49, 0xC9, 0x29, 0xA9, 0x69, 0xE9, 0x19, 0x99, 0x59, 0xD9, 0x39, 0xB9, 0x79, 0xF9, 
  0x05, 0x85, 0x45, 0xC5, 0x25, 0xA5, 0x65, 0xE5, 0x15, 0x95, 0x55, 0xD5, 0x35, 0xB5, 0x75, 0xF5,
  0x0D, 0x8D, 0x4D, 0xCD, 0x2D, 0xAD, 0x6D, 0xED, 0x1D, 0x9D, 0x5D, 0xDD, 0x3D, 0xBD, 0x7D, 0xFD,
  0x03, 0x83, 0x43, 0xC3, 0x23, 0xA3, 0x63, 0xE3, 0x13, 0x93, 0x53, 0xD3, 0x33, 0xB3, 0x73, 0xF3, 
  0x0B, 0x8B, 0x4B, 0xCB, 0x2B, 0xAB, 0x6B, 0xEB, 0x1B, 0x9B, 0x5B, 0xDB, 0x3B, 0xBB, 0x7B, 0xFB,
  0x07, 0x87, 0x47, 0xC7, 0x27, 0xA7, 0x67, 0xE7, 0x17, 0x97, 0x57, 0xD7, 0x37, 0xB7, 0x77, 0xF7, 
  0x0F, 0x8F, 0x4F, 0xCF, 0x2F, 0xAF, 0x6F, 0xEF, 0x1F, 0x9F, 0x5F, 0xDF, 0x3F, 0xBF, 0x7F, 0xFF
};

unsigned int bit_reversal(const unsigned int i){
  return (BitReverseTable256[i & 0xFF] << 24) | 
    (BitReverseTable256[(i >> 8) & 0xFF] << 16) | 
    (BitReverseTable256[(i >> 16) & 0xFF] << 8) |
    (BitReverseTable256[(i >> 24) & 0xFF]);
}

Count the number of bits

Hamming weight is faster than lookup table according to Google’s “Director of Engineering” Hiring Test. The actual performance test is here.
const uint64_t m1  = 0x55555555; // 0101...
const uint64_t m2  = 0x33333333; // 00110011...
const uint64_t m4  = 0x0F0F0F0F; // 0000111100001111...
const uint64_t m8  = 0x00FF00FF; // 8 zeros, 8 ones...
const uint64_t m16 = 0x0000FFFF; // 16 zeros, 16 ones...
const uint64_t m32 = 0x00000000ffffffff; // 32 zeros, 32 ones
const uint64_t h01 = 0x0101010101010101; //the sum of 256 to the power of 0,1,2,3...

//This is a naive implementation, shown for comparison,
//and to help in understanding the better functions.
//This algorithm uses 24 arithmetic operations (shift, add, and).
int popcount64a(uint64_t x)
{
    x = (x & m1 ) + ((x >>  1) & m1 ); //put count of each  2 bits into those  2 bits 
    x = (x & m2 ) + ((x >>  2) & m2 ); //put count of each  4 bits into those  4 bits 
    x = (x & m4 ) + ((x >>  4) & m4 ); //put count of each  8 bits into those  8 bits 
    x = (x & m8 ) + ((x >>  8) & m8 ); //put count of each 16 bits into those 16 bits 
    x = (x & m16) + ((x >> 16) & m16); //put count of each 32 bits into those 32 bits 
    x = (x & m32) + ((x >> 32) & m32); //put count of each 64 bits into those 64 bits 
    return x;
}

//This uses fewer arithmetic operations than any other known  
//implementation on machines with slow multiplication.
//This algorithm uses 17 arithmetic operations.
int popcount64b(uint64_t x)
{
    x -= (x >> 1) & m1;             //put count of each 2 bits into those 2 bits
    x = (x & m2) + ((x >> 2) & m2); //put count of each 4 bits into those 4 bits 
    x = (x + (x >> 4)) & m4;        //put count of each 8 bits into those 8 bits 
    x += x >>  8;  //put count of each 16 bits into their lowest 8 bits
    x += x >> 16;  //put count of each 32 bits into their lowest 8 bits
    x += x >> 32;  //put count of each 64 bits into their lowest 8 bits
    return x & 0x7f;
}

//This uses fewer arithmetic operations than any other known  
//implementation on machines with fast multiplication.
//This algorithm uses 12 arithmetic operations, one of which is a multiply.
int popcount64c(uint64_t x)
{
    x -= (x >> 1) & m1;             //put count of each 2 bits into those 2 bits
    x = (x & m2) + ((x >> 2) & m2); //put count of each 4 bits into those 4 bits 
    x = (x + (x >> 4)) & m4;        //put count of each 8 bits into those 8 bits 
    return (x * h01) >> 56;  //returns left 8 bits of x + (x<<8) + (x<<16) + (x<<24) + ... 
}