Revision 4.2
Tim PenheyEach style point has a summary for which additional information is available by toggling the accompanying arrow button that looks this way: ▶. You may toggle all summaries with the big arrow button:
Hooray! Now you know you can expand points to get more details. Alternatively, there's an "expand all" at the top of this document.
As every C++ programmer knows, the language has many powerful features, but this power brings with it complexity, which in turn can make code more bug-prone and harder to read and maintain.
The goal of this guide is to manage this complexity by describing in detail the dos and don'ts of writing C++ code. These rules exist to keep the code base manageable while still allowing coders to use C++ language features productively.
Style, also known as readability, is what we call the conventions that govern our C++ code. The term Style is a bit of a misnomer, since these conventions cover far more than just source file formatting.
One way in which we keep the code base manageable is by enforcing consistency. It is very important that any programmer be able to look at another's code and quickly understand it. Maintaining a uniform style and following conventions means that we can more easily use "pattern-matching" to infer what various symbols are and what invariants are true about them. Creating common, required idioms and patterns makes code much easier to understand. In some cases there might be good arguments for changing certain style rules, but we nonetheless keep things as they are in order to preserve consistency.
Another issue this guide addresses is that of C++ feature bloat. C++ is a huge language with many advanced features. In some cases we constrain, or even ban, use of certain features. We do this to keep code simple and to avoid the various common errors and problems that these features can cause. This guide lists these features and explains why their use is restricted.
Note that this guide is not a C++ tutorial: we assume that the reader is familiar with the language.
In general, every .cpp file should have an associated
.h file. There are some common exceptions, such as
unit tests and small .cpp files containing just a
main() function.
Correct use of header files can make a huge difference to the readability, size and performance of your code.
The following rules will guide you through the various pitfalls of using header files.
#define guards to
prevent multiple inclusion. The format of the symbol name
should be
<PROJECT>_<PATH>_<FILE>_H_.
To guarantee uniqueness, they should be based on the full path
in a project's source tree. For example, the file
foo/src/bar/baz.h in project foo should
have the following guard:
#ifndef FOO_BAR_BAZ_H_ #define FOO_BAR_BAZ_H_ ... #endif // FOO_BAR_BAZ_H_
#include when a forward declaration
would suffice.
When you include a header file you introduce a dependency that will cause your code to be recompiled whenever the header file changes. If your header file includes other header files, any change to those files will cause any code that includes your header to be recompiled. Therefore, we prefer to minimize includes, particularly includes of header files in other header files.
You can significantly reduce the number of header files you
need to include in your own header files by using forward
declarations. For example, if your header file uses the
File class in ways that do not require access to
the declaration of the File class, your header
file can just forward declare class File; instead
of having to #include "file/base/file.h".
How can we use a class Foo in a header file
without access to its definition?
Foo* or
Foo&.
Foo. (One
exception is if an argument Foo
or Foo const& has a
non-explicit, one-argument constructor,
in which case we need the full definition to support
automatic type conversion.)
Foo. This is because static data members
are defined outside the class definition.
On the other hand, you must include the header file for
Foo if your class subclasses Foo or
has a data member of type Foo.
Sometimes it makes sense to have pointer (or better,
unique_ptr)
members instead of object members. However, this complicates code
readability and imposes a performance penalty, so avoid doing
this transformation if the only purpose is to minimize includes
in header files.
Of course, .cpp files typically do require the
definitions of the classes they use, and usually have to
include several header files.
Note:
If you use a symbol Foo in your source file, you
should bring in a definition for Foo yourself,
either via an #include or via a forward declaration. Do not
depend on the symbol being brought in transitively via headers
not directly included. One exception is if Foo
is used in myfile.cpp, it's ok to #include (or
forward-declare) Foo in myfile.h,
instead of myfile.cpp.
Definition: You can declare functions in a way that allows the compiler to expand them inline rather than calling them through the usual function call mechanism.
Pros: Inlining a function can generate more efficient object code, as long as the inlined function is small. Feel free to inline accessors and mutators, and other short, performance-critical functions.
Cons: Overuse of inlining can actually make programs slower. Depending on a function's size, inlining it can cause the code size to increase or decrease. Inlining a very small accessor function will usually decrease code size while inlining a very large function can dramatically increase code size. On modern processors smaller code usually runs faster due to better use of the instruction cache.
Decision:
A decent rule of thumb is to not inline a function if it is more than 10 lines long. Beware of destructors, which are often longer than they appear because of implicit member- and base-destructor calls!
Another useful rule of thumb: it's typically not cost effective to inline functions with loops or switch statements (unless, in the common case, the loop or switch statement is never executed).
It is important to know that functions are not always inlined even if they are declared as such; for example, virtual and recursive functions are not normally inlined. Usually recursive functions should not be inline. The main reason for making a virtual function inline is to place its definition in the class, either for convenience or to document its behavior, e.g., for accessors and mutators.
-inl.h suffix to define
complex inline functions when needed.
The definition of an inline function needs to be in a header
file, so that the compiler has the definition available for
inlining at the call sites. However, implementation code
properly belongs in .cpp files, and we do not like
to have much actual code in .h files unless there
is a readability or performance advantage.
If an inline function definition is short, with very little,
if any, logic in it, you should put the code in your
.h file. For example, accessors and mutators
should certainly be inside a class definition. More complex
inline functions may also be put in a .h file for
the convenience of the implementer and callers, though if this
makes the .h file too unwieldy you can instead
put that code in a separate -inl.h file.
This separates the implementation from the class definition,
while still allowing the implementation to be included where
necessary.
Another use of -inl.h files is for definitions of
function templates. This can be used to keep your template
definitions easy to read.
Do not forget that a -inl.h file requires a
#define guard just
like any other header file.
Parameters to C/C++ functions are either input to the
function, output from the function, or both. Input parameters
are usually values or const references, while output
and input/output parameters will be non-const
references or pointers to non-const. When ordering function
parameters, put all output parameters before any input-only parameters.
In particular, do not add new parameters to the end of the function just
because they are new; place new output parameters before the input-only
parameters.
This is not a hard-and-fast rule. Parameters that are both input and output (often classes/structs) muddy the waters, and, as always, consistency with related functions may require you to bend the rule.
.h, your
project's private
.h, other libraries' .h, .C library, C++ library,
All of a project's header files should be
listed as descendants of the project's source directory
without use of UNIX directory shortcuts . (the current
directory) or .. (the parent directory). For
example,
my-awesome-project/src/base/logging.h
should be included as
#include "base/logging.h"
In dir/foo.cpp or dir/foo_test.cpp,
whose main purpose is to implement or test the stuff in
dir2/foo2.h, order your includes as
follows:
dir2/foo2.h (preferred location
— see details below)..h files.
.h files.
.h files.
The preferred ordering reduces hidden dependencies. We want
every header file to be compilable on its own. The easiest
way to achieve this is to make sure that every one of them is
the first .h file #included in some
.cpp.
dir/foo.cpp and
dir2/foo2.h are often in the same
directory (e.g. base/test_basictypes.cpp and
base/basictypes.h), but can be in different
directories too.
Within each section it is nice to order the includes alphabetically.
For example, the includes in
my-awesome-project/src/foo/internal/fooserver.cpp
might look like this:
#include "foo/public/fooserver.h" // Preferred location.
#include "base/basictypes.h"
#include "base/commandlineflags.h"
#include "foo/public/bar.h"
#include <sys/types.h>
#include <unistd.h>
#include <hash_map>
#include <vector>
.cpp files are encouraged. With
named namespaces, choose the name based on the
project, and possibly its path.
Do not use a using-directive in a header file.
Definition: Namespaces subdivide the global scope into distinct, named scopes, and so are useful for preventing name collisions in the global scope.
Pros:
Namespaces provide a (hierarchical) axis of naming, in addition to the (also hierarchical) name axis provided by classes.
For example, if two different projects have a class
Foo in the global scope, these symbols may
collide at compile time or at runtime. If each project
places their code in a namespace, project1::Foo
and project2::Foo are now distinct symbols that
do not collide.
Cons:
Namespaces can be confusing, because they provide an additional (hierarchical) axis of naming, in addition to the (also hierarchical) name axis provided by classes.
Use of unnamed spaces in header files can easily cause violations of the C++ One Definition Rule (ODR).
Decision:
Use namespaces according to the policy described below.
Unnamed Namespaces
.cpp files, to avoid runtime naming
conflicts:
namespace // This is in a .cpp file.
{
// The content of a namespace is not indented
enum { UNUSED, EOF, ERROR }; // Commonly used tokens.
bool AtEof() { return pos_ == EOF; } // Uses our namespace's EOF.
} // namespace
However, file-scope declarations that are
associated with a particular class may be declared
in that class as types, static data members or
static member functions rather than as members of
an unnamed namespace. Terminate the unnamed
namespace as shown, with a comment //
namespace.
.h
files.
Named Namespaces
Named namespaces should be used as follows:
// In the .h file
namespace mynamespace
{
// All declarations are within the namespace scope.
// Notice the lack of indentation.
class MyClass
{
public:
...
void foo();
};
} // namespace mynamespace// In the .cpp file
namespace mynamespace
{
// Definition of functions is within scope of the namespace.
void MyClass::foo()
{
...
}
} // namespace mynamespace
The typical .cpp file might have more
complex detail, including the need to reference classes
in other namespaces.
#include "a.h"
DEFINE_BOOL(someflag, false, "dummy flag");
class C; // Forward declaration of class C in the global namespace.
namespace a { class A; } // Forward declaration of a::A.
namespace b
{
...code for b... // Code goes against the left margin.
} // namespace bstd, not even forward declarations of
standard library classes. Declaring entities in
namespace std is undefined behavior,
i.e., not portable. To declare entities from the
standard library, include the appropriate header
file.
.cpp file, and in functions,
methods or classes in .h files.
// OK in .cpp files. // Must be in a function, method or class in .h files. using ::foo::bar;
.cpp file, anywhere inside the named
namespace that wraps an entire .h file,
and in functions and methods.
// Shorten access to some commonly used names in .cpp files.
namespace fbz = ::foo::bar::baz;
// Shorten access to some commonly used names (in a .h file).
namespace librarian
{
// The following alias is available to all files including
// this header (in namespace librarian):
// alias names should therefore be chosen consistently
// within a project.
namespace pd_s = ::pipeline_diagnostics::sidetable;
inline void my_inline_function()
{
// namespace alias local to a function (or method).
namespace fbz = ::foo::bar::baz;
...
}
} // namespace librarianNote that an alias in a .h file is visible to everyone #including that file, so public headers (those available outside a project) and headers transitively #included by them, should avoid defining aliases, as part of the general goal of keeping public APIs as small as possible.
Definition: A class can define another class within it; this is also called a member class.
class Foo
{
private:
// Bar is a member class, nested within Foo.
class Bar
{
...
};
};
Pros:
This is useful when the nested (or member) class is only used
by the enclosing class; making it a member puts it in the
enclosing class scope rather than polluting the outer scope
with the class name. Nested classes can be forward declared
within the enclosing class and then defined in the
.cpp file to avoid including the nested class
definition in the enclosing class declaration, since the
nested class definition is usually only relevant to the
implementation.
Cons:
Nested classes can be forward-declared only within the
definition of the enclosing class. Thus, any header file
manipulating a Foo::Bar* pointer will have to
include the full class declaration for Foo.
Decision: Do not make nested classes public unless they are actually part of the interface, e.g., a class that holds a set of options for some method.
Pros: Nonmember and static member functions can be useful in some situations. Putting nonmember functions in a namespace avoids polluting the global namespace.
Cons: Nonmember and static member functions may make more sense as members of a new class, especially if they access external resources or have significant dependencies.
Decision:
Sometimes it is useful, or even necessary, to define a function not bound to a class instance. Such a function can be either a static member or a nonmember function. Nonmember functions should not depend on external variables, and should nearly always exist in a namespace. Rather than creating classes only to group static member functions which do not share static data, use namespaces instead.
Functions defined in the same compilation unit as production classes may introduce unnecessary coupling and link-time dependencies when directly called from other compilation units; static member functions are particularly susceptible to this. Consider extracting a new class, or placing the functions in a namespace possibly in a separate library.
If you must define a nonmember function and it is only
needed in its .cpp file, use an unnamed
namespace or static
linkage (eg static int foo() {...}) to limit
its scope.
C++ allows you to declare variables anywhere in a function. We encourage you to declare them in as local a scope as possible, and as close to the first use as possible. This makes it easier for the reader to find the declaration and see what type the variable is and what it was initialized to. In particular, initialization should be used instead of declaration and assignment, e.g.
int i; i = f(); // Bad -- initialization separate from declaration.
int j = g(); // Good -- declaration has initialization.
Note that gcc implements for (int i = 0; i
< 10; ++i) correctly (the scope of i is
only the scope of the for loop), so you can then
reuse i in another for loop in the
same scope. It also correctly scopes declarations in
if and while statements, e.g.
while (char const* p = strchr(str, '/')) str = p + 1;
There is one caveat: if the variable is an object, its constructor is invoked every time it enters scope and is created, and its destructor is invoked every time it goes out of scope.
// Inefficient implementation:
for (int i = 0; i < 1000000; ++i)
{
Foo f; // My ctor and dtor get called 1000000 times each.
f.do_something(i);
}It may be more efficient to declare such a variable used in a loop outside that loop:
Foo f; // My ctor and dtor get called once each.
for (int i = 0; i < 1000000; ++i)
{
f.do_something(i);
}Definition: It is possible to perform initialization in the body of the constructor.
Pros: Convenience in typing. No need to worry about whether the class has been initialized or not.
Cons: The problems with doing work in constructors are:
main(), possibly breaking some implicit
assumptions in the constructor code.
Decision: Constructors should not make virtual calls to functions, access potentially uninitialized global variables, etc.
Definition:
The default constructor is called when we create a
class object with no arguments. It is always called when
calling new[] (for arrays).
Pros: Initializing structures by default makes debugging much easier.
Cons: Extra work for you, the code writer.
Decision:
If your class defines POD member variables and has no other constructors you must define a default constructor (one that takes no arguments). It should initialize the object in such a way that its internal state is consistent and valid.
The reason for this is that if you have no other constructors and do not define a default constructor, the compiler will generate one for you. This compiler generated constructor may not initialize your object sensibly.
If your class is composed from and/or inherits from an existing class or classes but you add no new member variables, you are not required to have a default constructor.
If your class has value semantics then consider making the
class invariants such that the default constructor is cheap.
For example, initialising member pointers to nullptr
and allocating on first use.
explicit for constructors with
one argument.
Definition:
Normally, if a constructor takes one argument, it can be used
as a conversion. For instance, if you define
Foo::Foo(string name) and then pass a string to a
function that expects a Foo, the constructor will
be called to convert the string into a Foo and
will pass the Foo to your function for you. This
can be convenient but is also a source of trouble when things
get converted and new objects created without you meaning them
to. Declaring a constructor explicit prevents it
from being invoked implicitly as a conversion.
Pros: Avoids undesirable conversions.
Cons: Avoids desirable conversions.
Decision:
We require all single argument constructors to be
explicit. Always put explicit in front of
one-argument constructors in the class definition:
explicit Foo(string name);
The exception is copy constructors, which, in the rare
cases when we allow them, should probably not be
explicit.
Classes that are intended to be
transparent wrappers around other classes are also
exceptions.
Such exceptions should be clearly marked with comments.
= delete;.
Definition: The copy constructor and assignment operator are used to create copies of objects. The copy constructor is implicitly invoked by the compiler in some situations, e.g. passing objects by value.
Pros:
Copy constructors make it easy to copy objects. STL
containers require that all contents be copyable and
assignable. Copy constructors can be more efficient than
CopyFrom()-style workarounds because they combine
construction with copying, the compiler can elide them in some
contexts, and they make it easier to avoid heap allocation.
Cons: Implicit copying of objects in C++ is a rich source of bugs and of performance problems. It also reduces readability, as it becomes hard to track which objects are being passed around by value as opposed to by reference, and therefore where changes to an object are reflected.
Decision:
Few classes need to be copyable. Most should have neither a copy constructor nor an assignment operator. In many situations, a pointer or reference will work just as well as a copied value, with better performance. For example, you can pass function parameters by reference or pointer instead of by value, and you can store pointers rather than objects in an STL container.
If your class needs to be copyable, prefer providing a copy method,
such as CopyFrom() or Clone(), rather than
a copy constructor, because such methods cannot be invoked
implicitly. If a copy method is insufficient in your situation
(e.g. for performance reasons, or because your class needs to be
stored by value in an STL container), provide both a copy
constructor and assignment operator.
If your class does not need a copy constructor or assignment operator, you must explicitly disable them.
struct only for passive objects that carry data;
everything else is a class.
The struct and class keywords behave
almost identically in C++. We add our own semantic meanings
to each keyword, so you should use the appropriate keyword for
the data-type you're defining.
structs should be used for passive objects that carry
data, and may have associated constants, but lack any functionality
other than access/setting the data members. The
accessing/setting of fields is done by directly accessing the
fields rather than through method invocations. Methods should
not provide behavior but should only be used to set up the
data members, e.g., constructor, destructor,
initialize(), reset(),
validate().
If more functionality is required, a class is more
appropriate. If in doubt, make it a class.
For consistency with STL, you can use struct
instead of class for functors and traits.
public.
Definition: When a sub-class inherits from a base class, it includes the definitions of all the data and operations that the parent base class defines. In practice, inheritance is used in two major ways in C++: implementation inheritance, in which actual code is inherited by the child, and interface inheritance, in which only method names are inherited.
Pros: Implementation inheritance reduces code size by re-using the base class code as it specializes an existing type. Because inheritance is a compile-time declaration, you and the compiler can understand the operation and detect errors. Interface inheritance can be used to programmatically enforce that a class expose a particular API. Again, the compiler can detect errors, in this case, when a class does not define a necessary method of the API.
Cons: For implementation inheritance, because the code implementing a sub-class is spread between the base and the sub-class, it can be more difficult to understand an implementation. The sub-class cannot override functions that are not virtual, so the sub-class cannot change implementation. The base class may also define some data members, so that specifies physical layout of the base class.
Decision:
All inheritance should be public. If you want to
do private inheritance, you should be including an instance of
the base class as a member instead.
Do not overuse implementation inheritance. Composition is
often more appropriate. Try to restrict use of inheritance
to the "is-a" case: Bar subclasses
Foo if it can reasonably be said that
Bar "is a kind of" Foo.
Make your destructor virtual if necessary. If
your class has virtual methods, its destructor
should be virtual.
Limit the use of protected to those member
functions that might need to be accessed from subclasses.
Note that data members should
be private.
When redefining an inherited virtual method (both pure
and non-pure), explicitly declare it override
in the declaration of the derived class. Rationale: using
override allows the compiler to consistently
detect attempts to override methods that have been changed
or completely removed. It also makes it straightforward for
a reader to determine if a method is virtual or not.
Definition: Multiple inheritance allows a sub-class to have more than one base class. We distinguish between base classes that are interfaces and those that have an implementation.
Pros: Multiple implementation inheritance may let you re-use even more code than single inheritance (see Inheritance).
Cons: Only very rarely is multiple implementation inheritance actually useful. When multiple implementation inheritance seems like the solution, you can usually find a different, more explicit, and cleaner solution.
Decision: Multiple inheritance is allowed only when all superclasses, with the possible exception of the first one, are interfaces.
Definition:
A class is an interface if it meets the following requirements:
= 0") methods
and static methods (but see below for destructor).
An interface class can never be directly instantiated because of the pure virtual method(s) it declares. To make sure all implementations of the interface can be destroyed correctly, they must also declare a virtual, or protected, destructor (in an exception to the first rule, this should not be pure). See Stroustrup, The C++ Programming Language, 3rd edition, section 12.4 for details.
Definition:
A class can define that operators such as + and
/ operate on the class as if it were a built-in
type.
Pros:
Can make code appear more intuitive because a class will
behave in the same way as built-in types (such as
int). Overloaded operators are more playful
names for functions that are less-colorfully named, such as
equals() or add(). For some
template functions to work correctly, you may need to define
operators.
Cons: While operator overloading can make code more intuitive, it has several drawbacks:
Equals() is much
easier than searching for relevant invocations of
==.
Foo + 4 may do one thing,
while &Foo + 4 does something totally
different. The compiler does not complain for either of
these, making this very hard to debug.
operator&, it
cannot safely be forward-declared.
Decision:
In general, do overload operators where appropriate.
See also Copy Constructors and Function Overloading.
private, and provide
access to them through accessor functions as needed (for
technical reasons, we allow data members of a test fixture class
to be protected when using
Google Test). Typically a variable would be
called foo and the accessor function
get_foo(). You may also want a mutator function
set_foo().
Exception: static const data members need not
be private.
The definitions of accessors are usually inlined in the header file.
See also Inheritance and Function Names.
public: before private:, methods
before data members (variables), etc.
Your class definition should start with its public:
section, followed by its protected: section and
then its private: section. If any of these sections
are empty, omit them.
Within each section, the declarations generally should be in the following order:
static const data members)static const data members)
Friend declarations should always be in the private section, and
the DISALLOW_COPY_AND_ASSIGN macro invocation
should be at the end of the private: section. It
should be the last thing in the class. See Copy Constructors.
Method definitions in the corresponding .cpp file
should be the same as the declaration order, as much as possible.
Do not put large method definitions inline in the class definition. Usually, only trivial or performance-critical, and very short, methods may be defined inline. See Inline Functions for more details.
We recognize that long functions are sometimes appropriate, so no hard limit is placed on functions length. If a function exceeds about 40 lines, think about whether it can be broken up without harming the structure of the program.
Even if your long function works perfectly now, someone modifying it in a few months may add new behavior. This could result in bugs that are hard to find. Keeping your functions short and simple makes it easier for other people to read and modify your code.
You could find long and complicated functions when working with some code. Do not be intimidated by modifying existing code: if working with such a function proves to be difficult, you find that errors are hard to debug, or you want to use a piece of it in several different contexts, consider breaking up the function into smaller and more manageable pieces.
const.
Definition:
In C, if a function needs to modify a variable, the
parameter must use a pointer, eg int foo(int
*pval). In C++, the function can alternatively
declare a reference parameter: int foo(int& val).
Pros:
Defining a parameter as reference avoids ugly code like
(*pval)++. Necessary for some applications like
copy constructors. Makes it clear, unlike with pointers, that
NULL is not a possible value.
Cons: References can be confusing, as they have value syntax but pointer semantics.
Decision:
Within function parameter lists all references must be
const:
void foo(string const& in, string* out);
In fact it is a very strong convention in Unity code that input
arguments are values or const references while output
arguments are pointers. Input parameters may be const
pointers. Non-const reference parameters are allowed
but there must be a valid reason for it, a strong preference is
given to const reference parameters.
One case when you might want an input parameter to be a
const pointer is if you want to emphasize that the
argument is not copied, so it must exist for the lifetime of the
object; it is usually best to document this in comments as
well. STL adapters such as bind2nd and
mem_fun do not permit reference parameters, so
you must declare functions with pointer parameters in these
cases, too.
Definition:
You may write a function that takes a
string const& and overload it with another that
takes char const*.
class MyClass {
public:
void analyze(string const& text);
void analyze(char const* text, size_t textlen);
};Pros: Overloading can make code more intuitive by allowing an identically-named function to take different arguments. It may be necessary for templatized code, and it can be convenient for Visitors.
Cons: If a function is overloaded by the argument types alone, a reader may have to understand C++'s complex matching rules in order to tell what's going on. Also many people are confused by the semantics of inheritance if a derived class overrides only some of the variants of a function.
Decision:
If you want to overload a function, consider qualifying the
name with some information about the arguments, e.g.,
append_string(), AppendInt() rather
than just append().
Pros: Often you have a function that uses lots of default values, but occasionally you want to override the defaults. Default parameters allow an easy way to do this without having to define many functions for the rare exceptions.
Cons: People often figure out how to use an API by looking at existing code that uses it. Default parameters are more difficult to maintain because copy-and-paste from previous code may not reveal all the parameters. Copy-and-pasting of code segments can cause major problems when the default arguments are not appropriate for the new code.
Decision:
Except as described below, we require all arguments to be explicitly specified, to force programmers to consider the API and the values they are passing for each argument rather than silently accepting defaults they may not be aware of.
One specific exception is when default arguments are used to simulate variable-length argument lists.
// Support up to 4 params by using a default empty AlphaNum.
string str_cat(AlphaNum const& a,
AlphaNum const& b = gEmptyAlphaNum,
AlphaNum const& c = gEmptyAlphaNum,
AlphaNum const& d = gEmptyAlphaNum);alloca().
Pros:
Variable-length arrays have natural-looking syntax. Both
variable-length arrays and alloca() are very
efficient.
Cons: Variable-length arrays and alloca are not part of Standard C++. More importantly, they allocate a data-dependent amount of stack space that can trigger difficult-to-find memory overwriting bugs: "It ran fine on my machine, but dies mysteriously in production".
Decision:
Use a safe allocator instead, such as
unique_ptr.
friend classes and functions,
within reason.
Friends should usually be defined in the same file so that the
reader does not have to look in another file to find uses of
the private members of a class. A common use of
friend is to have a FooBuilder class
be a friend of Foo so that it can construct the
inner state of Foo correctly, without exposing
this state to the world. In some cases it may be useful to
make a unit test class a friend of the class it tests.
Friends extend, but do not break, the encapsulation boundary of a class. In some cases this is better than making a member public when you want to give only one other class access to it. However, most classes should interact with other classes solely through their public members.
static_cast<>(). Do not use
other cast formats like int y = (int)x; or
int y = int(x);.
Definition: C++ introduced a different cast system from C that distinguishes the types of cast operations.
Pros:
The problem with C casts is the ambiguity of the operation;
sometimes you are doing a conversion (e.g.,
(int)3.5) and sometimes you are doing a
cast (e.g., (int)"hello"); C++ casts
avoid this. Additionally C++ casts are more visible when
searching for them.
Cons: The syntax is nasty.
Decision:
Do not use C-style casts. Instead, use these C++-style casts.
static_cast as the equivalent of a
C-style cast that does value conversion, or when you need to explicitly up-cast
a pointer from a class to its superclass.
const_cast to remove the const
qualifier (see const).
reinterpret_cast to do unsafe
conversions of pointer types to and from integer and
other pointer types. Use this only if you know what you are
doing and you understand the aliasing issues.
dynamic_cast except in test code.
If you need to know type information at runtime in this way
outside of a unit test, you probably have a design
flaw.
Definition:
Streams are a replacement for printf() and
scanf().
Pros:
With streams, you do not need to know the type of the object
you are printing. You do not have problems with format
strings not matching the argument list. (Though with gcc, you
do not have that problem with printf either.) Streams
have automatic constructors and destructors that open and close the
relevant files.
Cons:
Streams make it difficult to do functionality like
pread(). Some formatting (particularly the common
format string idiom %.*s) is difficult if not
impossible to do efficiently using streams without using
printf-like hacks. Streams do not support operator
reordering (the %1s directive), which is helpful for
internationalization.
Decision:
Do not use streams, except where required by a logging interface.
Use printf-like routines instead.
There are various pros and cons to using streams, but in this case, as in many other cases, consistency trumps the debate. Do not use streams in your code.
Extended Discussion
There has been debate on this issue, so this explains the
reasoning in greater depth. Recall the Only One Way
guiding principle: we want to make sure that whenever we
do a certain type of I/O, the code looks the same in all
those places. Because of this, we do not want to allow
users to decide between using streams or using
printf plus Read/Write/etc. Instead, we should
settle on one or the other. We made an exception for logging
because it is a pretty specialized application, and for
historical reasons.
Proponents of streams have argued that streams are the obvious choice of the two, but the issue is not actually so clear. For every advantage of streams they point out, there is an equivalent disadvantage. The biggest advantage is that you do not need to know the type of the object to be printing. This is a fair point. But, there is a downside: you can easily use the wrong type, and the compiler will not warn you. It is easy to make this kind of mistake without knowing when using streams.
cout << this; // Prints the address cout << *this; // Prints the contents
The compiler does not generate an error because
<< has been overloaded. We discourage
overloading for just this reason.
Some say printf formatting is ugly and hard to
read, but streams are often no better. Consider the following
two fragments, both with the same typo. Which is easier to
discover?
cerr << "Error connecting to '" << foo->bar()->hostname.first
<< ":" << foo->bar()->hostname.second << ": " << strerror(errno);
fprintf(stderr, "Error connecting to '%s:%u: %s",
foo->bar()->hostname.first, foo->bar()->hostname.second,
strerror(errno));And so on and so forth for any issue you might bring up. (You could argue, "Things would be better with the right wrappers," but if it is true for one scheme, is it not also true for the other? Also, remember the goal is to make the language smaller, not add yet more machinery that someone has to learn.)
Either path would yield different advantages and
disadvantages, and there is not a clearly superior
solution. The simplicity doctrine mandates we settle on
one of them though, and the majority decision was on
printf + read/write.
++i) of the increment and
decrement operators with iterators and other template objects.
Definition:
When a variable is incremented (++i or
i++) or decremented (--i or
i--) and the value of the expression is not used,
one must decide whether to preincrement (decrement) or
postincrement (decrement).
Pros:
When the return value is ignored, the "pre" form
(++i) is never less efficient than the "post"
form (i++), and is often more efficient. This is
because post-increment (or decrement) requires a copy of
i to be made, which is the value of the
expression. If i is an iterator or other
non-scalar type, copying i could be expensive.
Since the two types of increment behave the same when the
value is ignored, why not just always pre-increment?
Cons:
The tradition developed, in C, of using post-increment when
the expression value is not used, especially in for
loops. Some find post-increment easier to read, since the
"subject" (i) precedes the "verb" (++),
just like in English.
Decision: For simple scalar (non-object) values there is no reason to prefer one form and we allow either. For iterators and other template types, use pre-increment.
const whenever
it makes sense to do so.
Definition:
Declared variables and parameters can be preceded by the
keyword const to indicate the variables are not
changed (e.g., int const foo). Class functions
can have the const qualifier to indicate the
function does not change the state of the class member
variables (e.g., class Foo { int bar(char c) const;
};).
Pros: Easier for people to understand how variables are being used. Allows the compiler to do better type checking, and, conceivably, generate better code. Helps people convince themselves of program correctness because they know the functions they call are limited in how they can modify your variables. Helps people know what functions are safe to use without locks in multi-threaded programs.
Cons:
const is viral: if you pass a const
variable to a function, that function must have const
in its prototype (or the variable will need a
const_cast). This can be a particular problem
when calling library functions.
Decision:
const variables, data members, methods and
arguments add a level of compile-time type checking; it
is better to detect errors as soon as possible.
Therefore we strongly recommend that you use
const whenever it makes sense to do so:
const.
const whenever
possible. Accessors should almost always be
const. Other methods should be const if they do
not modify any data members, do not call any
non-const methods, and do not return a
non-const pointer or non-const
reference to a data member.
const
whenever they do not need to be modified after
construction.
However, do not go crazy with const. Something like
int const* const* const x; is likely
overkill, even if it accurately describes how const x is.
Focus on what's really useful to know: in this case,
int const** x is probably sufficient.
The mutable keyword is allowed but is unsafe
when used with threads, so thread safety should be carefully
considered first.
Where to put the const
We favor the form int const* foo to
const int* foo. This keeps the const
with the type modifier (& or *).
That said, while we encourage putting const after the type,
we do not require it. But be consistent with the code around
you!
size_t where
appropriate. If a program needs a variable of a different size,
use a precise-width integer type from
<cstdint>, such as int16_t.
Definition:
C++ does not specify the sizes of its integer types. Typically
people assume that short is 16 bits,
int is 32 bits, long is 32 bits and
long long is 64 bits.
Pros: Uniformity of declaration.
Cons: The sizes of integral types in C++ can vary based on compiler and architecture.
Decision:
<cstdint> defines
types like int16_t, uint32_t,
int64_t, etc.
You should always use those in preference to
short, unsigned long long and the
like, when you need a guarantee on the size of an integer.
When appropriate, you are welcome to use standard
types like size_t and ptrdiff_t.
For integers we know can be "big",
use
int64_t.
printf() specifiers for some types are
not cleanly portable between 32-bit and 64-bit
systems. C99 defines some portable format
specifiers. Unfortunately, MSVC 7.1 does not
understand some of these specifiers and the
standard is missing a few, so we have to define our
own ugly versions in some cases (in the style of the
standard include file inttypes.h):
// printf macros for size_t, in the style of inttypes.h
#ifdef _LP64
#define __PRIS_PREFIX "z"
#else
#define __PRIS_PREFIX
#endif
// Use these macros after a % in a printf format string
// to get correct 32/64 bit behavior, like this:
// size_t size = records.size();
// printf("%"PRIuS"\n", size);
#define PRIdS __PRIS_PREFIX "d"
#define PRIxS __PRIS_PREFIX "x"
#define PRIuS __PRIS_PREFIX "u"
#define PRIXS __PRIS_PREFIX "X"
#define PRIoS __PRIS_PREFIX "o"| Type | DO NOT use | DO use | Notes |
|---|---|---|---|
void* (or any pointer) |
%lx |
%p |
|
int64_t |
%qd,
%lld
|
%"PRId64" |
|
uint64_t |
%qu,
%llu,
%llx
|
%"PRIu64",
%"PRIx64"
|
|
size_t |
%u |
%"PRIuS",
%"PRIxS"
|
C99 specifies %zu
|
ptrdiff_t |
%d |
%"PRIdS" |
C99 specifies %zd
|
Note that the PRI* macros expand to independent
strings which are concatenated by the compiler. Hence
if you are using a non-constant format string, you
need to insert the value of the macro into the format,
rather than the name. It is still possible, as usual,
to include length specifiers, etc., after the
% when using the PRI*
macros. So, e.g. printf("x = %30"PRIuS"\n",
x) would expand on 32-bit Linux to
printf("x = %30" "u" "\n", x), which the
compiler will treat as printf("x = %30u\n",
x).
sizeof(void*) !=
sizeof(int). Use intptr_t if
you want a pointer-sized integer.
int64_t/uint64_t
member will by default end up being 8-byte aligned on a 64-bit
system. If you have such structures being shared on disk
between 32-bit and 64-bit code, you will need to ensure
that they are packed the same on both architectures.
Most compilers offer a way to alter
structure alignment. For gcc, you can use
__attribute__((packed)). MSVC offers
#pragma pack() and
__declspec(align()).
LL or ULL suffixes as
needed to create 64-bit constants. For example:
int64_t my_value = 0x123456789LL; uint64_t my_mask = 3ULL << 48;
#ifdef _LP64 to choose between
the code variants. (But please avoid this if
possible, and keep any such changes localized.)
const variables to macros.
Macros mean that the code you see is not the same as the code the compiler sees. This can introduce unexpected behavior, especially since macros have global scope.
Luckily, macros are not nearly as necessary in C++ as they are
in C. Instead of using a macro to inline performance-critical
code, use an inline function. Instead of using a macro to
store a constant, use a const variable. Instead of
using a macro to "abbreviate" a long variable name, use a
reference. Instead of using a macro to conditionally compile code
... well, don't do that at all (except, of course, for the
#define guards to prevent double inclusion of
header files). It makes testing much more difficult.
Macros can do things these other techniques cannot, and you do see them in the codebase, especially in the lower-level libraries. And some of their special features (like stringifying, concatenation, and so forth) are not available through the language proper. But before using a macro, consider carefully whether there's a non-macro way to achieve the same result.
The following usage pattern will avoid many problems with macros; if you use macros, follow it whenever possible:
.h file.
#define macros right before you use them,
and #undef them right after.
#undef an existing macro before
replacing it with your own; instead, pick a name that's
likely to be unique.
## to generate function/class/variable
names.
0 for integers, 0.0 for reals,
nullptr for pointers, and '\0' for chars.
Use 0 for integers and 0.0 for reals.
This is not controversial.
For pointers (address values), C++11 added the nullptr
construct. This allows the compiler to do additional checks, and is the
preferred NULL pointer value.
Use '\0' for chars.
This is the correct type and also makes code more readable.
sizeof(varname) instead of
sizeof(type) whenever possible.
Use sizeof(varname) because it will update
appropriately if the type of the variable changes.
sizeof(type) may make sense in some cases,
but should generally be avoided because it can fall out of sync if
the variable's type changes.
Struct data; memset(&data, 0, sizeof(data));
memset(&data, 0, sizeof(Struct));
Definition: C++11 is the current ISO C++ standard. It contains significant changes both to the language and libraries from the older standard.
The most important consistency rules are those that govern naming. The style of a name immediately informs us what sort of thing the named entity is: a type, a variable, a function, a constant, a macro, etc., without requiring us to search for the declaration of that entity. The pattern-matching engine in our brains relies a great deal on these naming rules.
Naming rules are pretty arbitrary, but we feel that consistency is more important than individual preferences in this area, so regardless of whether you find them sensible or not, the rules are the rules.
How to Name
Give as descriptive a name as possible, within reason. Do not worry about saving horizontal space as it is far more important to make your code immediately understandable by a new reader. Examples of well-chosen names:
int num_errors; // Good. int num_completed_connections; // Good.
Poorly-chosen names use ambiguous abbreviations or arbitrary characters that do not convey meaning:
int n; // Bad - meaningless. int nerr; // Bad - ambiguous abbreviation. int n_comp_conns; // Bad - ambiguous abbreviation.
Type and variable names should typically be nouns: e.g.,
FileOpener,
num_errors.
Function names should typically be imperative (that is they
should be commands): e.g., open_file(),
set_num_errors(). There is an exception for
accessors, which, described more completely in Function Names, should be named
the same as the variable they access.
Abbreviations
Do not use abbreviations unless they are extremely well known outside your project. For example:
// Good // These show proper names with no abbreviations. int num_dns_connections; // Most people know what "DNS" stands for. int price_count_reader; // OK, price count. Makes sense.
// Bad! // Abbreviations can be confusing or ambiguous outside a small group. int wgc_connections; // Only your group knows what this stands for. int pc_reader; // Lots of things can be abbreviated "pc".
Never abbreviate by leaving out letters:
int error_count; // Good.
int error_cnt; // Bad.
_) or dashes (-). Follow the
convention that your
project
uses. If there is no consistent local pattern to follow, prefer "_".
Examples of acceptable file names:
my_useful_class.cpp
my-useful-class.cpp
myusefulclass.cpp
test_myusefulclass.cpp // _unittest and _regtest are deprecated.
C++ files should end in .cpp and header files
should end in .h.
Do not use filenames that already exist
in /usr/include, such as db.h.
In general, make your filenames very specific. For example,
use http_server_logs.h rather
than logs.h. A very common case is to have a
pair of files called, e.g., foo_bar.h
and foo_bar.cpp, defining a class
called FooBar.
Inline functions must be in a .h file. If your
inline functions are very short, they should go directly into your
.h file. However, if your inline functions
include a lot of code, they may go into a third file that
ends in -inl.h. In a class with a lot of inline
code, your class could have three files:
url_table.h // The class declaration. url_table.cpp // The class definition. url_table-inl.h // Inline functions that include lots of code.
See also the section -inl.h Files
MyExcitingClass, MyExcitingEnum.
The names of all types — classes, structs, typedefs, and enums — have the same naming convention. Type names should start with a capital letter and have a capital letter for each new word. No underscores. For example:
// classes and structs class UrlTable ... class UrlTableTester ... struct UrlTableProperties ... // typedefs typedef hash_map<UrlTableProperties*, string> PropertiesMap; // enums enum UrlTableErrors ...
my_exciting_local_variable,
my_exciting_member_variable.
Common Variable names
For example:
string table_name; // OK - uses underscore. string tablename; // OK - all lowercase.
string tableName; // Bad - mixed case.
Class Data Members
Data members (als