[历史归档]本文原发布于 cstriker1407.info 个人博客内容为历史存档仅供参考。发布时间2018-05-01 标题Protobuf学习笔记分类编程 标签CC·protobufprotobuf简单笔记Protobuf是什么为什么要用编写proto文件代码解释生成C文件常用API编译生成可执行代码编码风格标量类型列表protobuf 文件级别优化以下参考《 https://www.cnblogs.com/autyinjing/p/6495103.html 》《 https://www.cnblogs.com/stephen-liu74/archive/2013/01/02/2841485.html 》《 http://www.cnblogs.com/stephen-liu74/archive/2013/01/04/2842533.html 》《 https://www.cnblogs.com/davad/p/4871010.html 》Protobuf是什么Google Protocol Buffer(简称 Protobuf)是一种轻便高效的结构化数据存储格式平台无关、语言无关、可扩展可用于通讯协议和数据存储等领域。为什么要用平台无关语言无关可扩展提供了友好的动态库使用简单解析速度快比对应的XML快约20-100倍序列化数据非常简洁、紧凑与XML相比其序列化之后的数据量约为1/3到1/10。怎么安装源码下载地址 https://github.com/google/protobuf安装依赖的库autoconf automake libtool curl make g unzip安装 $ ./autogen.sh $ ./configure $make$makecheck $sudomakeinstall编写proto文件首先需要一个proto文件其中定义了我们程序中需要处理的结构化数据// Filename: addressbook.proto syntaxproto2; package addressbook; import src/help.proto; //举例用编译时去掉 message Person { required string name 1; required int32 id 2; optional string email 3; enum PhoneType { MOBILE 0; HOME 1; WORK 2; } message PhoneNumber { required string number 1; optional PhoneType type 2 ; } repeated PhoneNumber phone 4; } message AddressBook { repeated Person person_info 1; }代码解释// Filename: addressbook.proto 这一行是注释语法类似于Csyntax“proto2”; 表明使用protobuf的编译器版本为v2目前最新的版本为v3package addressbook; 声明了一个包名用来防止不同的消息类型命名冲突类似于 namespaceimport “src/help.proto”; 导入了一个外部proto文件中的定义类似于C中的 include 。message 是Protobuf中的结构化数据类似于C中的类可以在其中定义需要处理的数据required string name 1; 声明了一个名为name数据类型为string的required字段字段的标识号为1protobuf一共有三个字段修饰符- required该值是必须要设置的- optional 该字段可以有0个或1个值不超过1个- repeated该字段可以重复任意多次包括0次类似于C中的list使用建议除非确定某个字段一定会被设值否则使用optional代替required。string 是一种标量类型protobuf的所有标量类型请参考文末的标量类型列表。name 是字段名1 是字段的标识号在消息定义中每个字段都有唯一的一个数字标识号这些标识号是用来在消息的二进制格式中识别各个字段的一旦开始使用就不能够再改变。标识号的范围在1 ~ 229 - 1其中[1900019999]为Protobuf预留不能使用。Person 内部声明了一个enum和一个message这类似于C中的类内声明Person外部的结构可以用 Person.PhoneType 的方式来使用PhoneType。当使用外部package中的结构时要使用 pkgName.msgName.typeName 的格式每两层之间使用’.来连接类似C中的::。optional PhoneType type 2 ; 为type字段指定了一个默认值当没有为type设值时其值为HOME。另外一个proto文件中可以声明多个message在编译的时候他们会被编译成为不同的类。// Filename: addressbook.proto 这一行是注释语法类似于Csyntax“proto2”; 表明使用protobuf的编译器版本为v2目前最新的版本为v3package addressbook; 声明了一个包名用来防止不同的消息类型命名冲突类似于 namespaceimport “src/help.proto”; 导入了一个外部proto文件中的定义类似于C中的 include 。message 是Protobuf中的结构化数据类似于C中的类可以在其中定义需要处理的数据required string name 1; 声明了一个名为name数据类型为string的required字段字段的标识号为1protobuf一共有三个字段修饰符required该值是必须要设置的optional 该字段可以有0个或1个值不超过1个repeated该字段可以重复任意多次包括0次类似于C中的list生成C文件protoc是proto文件的编译器目前可以将proto文件编译成C、Java、Python三种代码文件编译格式如下protoc -IS R C _ D ∗ ∗ ∗ ∗ I R − − c p p _ o u t SRC\_D****IR --cpp\_outSRC_D∗∗∗∗IR−−cpp_outDST_DIR /path/to/file.proto上面的命令会生成xxx.pb.h 和 xxx.pb.cc两个C文件。使用C文在编写一个main.cc文件#includeiostream#includeaddressbook.pb.hintmain(intargc,constchar*argv[]){addressbook::AddressBook person;addressbook::Person*piperson.add_person_info();pi-set_name(aut);pi-set_id(1219);std::coutbefore clear(), id pi-id()std::endl;pi-clear_id();std::coutafter clear(), id pi-id()std::endl;pi-set_id(1087);if(!pi-has_email())pi-set_email(autyinjing126.com);addressbook::Person::PhoneNumber*pnpi-add_phone();pn-set_number(021-8888-8888);pnpi-add_phone();pn-set_number(138-8888-8888);pn-set_type(addressbook::Person::MOBILE);uint32_tsizeperson.ByteSize();unsignedcharbyteArray;person.SerializeToArray(byteArray,size);addressbook::AddressBook help_person;help_person.ParseFromArray(byteArray,size);addressbook::Person help_pihelp_person.person_info(0);std::cout*****************************std::endl;std::coutid: help_pi.id()std::endl;std::coutname: help_pi.name()std::endl;std::coutemail: help_pi.email()std::endl;for(inti0;ihelp_pi.phone_size();i){autohelp_pnhelp_pi.mutable_phone(i);std::coutphone_type: help_pn-type()std::endl;std::coutphone_number: help_pn-number()std::endl;}std::cout*****************************std::endl;return0;}常用APIprotoc为message的每个required字段和optional字段都定义了以下几个函数不限于这几个TypeNamexxx()const;//获取字段的值boolhas_xxx();//判断是否设值voidset_xxx(constTypeName);//设值voidclear_xxx();//使其变为默认值为每个repeated字段定义了以下几个 TypeName*add_xxx();//增加结点TypeNamexxx(int)const;//获取指定序号的结点类似于C的[]运算符TypeName*mutable_xxx(int);//类似于上一个但是获取的是指针intxxx_size();//获取结点的数量另外下面几个是常用的序列化函数boolSerializeToOstream(std::ostream*output)const;//输出到输出流中boolSerializeToString(string*output)const;//输出到stringboolSerializeToArray(void*data,intsize)const;//输出到字节流与之对应的反序列化函数boolParseFromIstream(std::istream*input);//从输入流解析boolParseFromString(conststringdata);//从string解析boolParseFromArray(constvoid*data,intsize);//从字节流解析其他常用的函数1boolIsInitialized();//检查是否所有required字段都被设值2size_tByteSize()const;//获取二进制字节序的大小官方API文档地址https://developers.google.com/protocol-buffers/docs/reference/overview编译生成可执行代码编译格式和普通的C代码一样但是要加上 -lprotobuf -pthreadg main.cc xxx.pb.cc -I $INCLUDE_PATH -L $LIB_PATH -lprotobuf -pthread输出结果before clear(), id 1219 after clear(), id 0 ***************************** id: 1087 name: aut email: autyinjing126.com phone_type: 1 phone_number: 021-8888-8888 phone_type: 0 phone_number: 138-8888-8888 *****************************怎么编码的protobuf之所以小且快就是因为使用变长的编码规则只保存有用的信息节省了大量空间。Base-128变长编码每个字节使用低7位表示数字除了最后一个字节其他字节的最高位都设置为1采用Little-Endian字节序。-数字10000 0001-数字3001010 1100 0000 0010000 0010 010 1100- 000 0010 010 1100- 100101100- 256 32 8 4 300ZigZag编码Base-128变长编码会去掉整数前面那些没用的0只保留低位的有效位然而负数的补码表示有很多的1所以protobuf先用ZigZag编码将所有的数值映射为无符号数然后使用Base-128编码ZigZag的编码规则如下(n 1) ^ (n 31) or (n 1) ^ (n 63)负数右移后高位全变成1再与左移一位后的值进行异或就把高位那些无用的1全部变成0了巧妙消息格式每一个Protocol Buffers的Message包含一系列的字段key/value每个字段由字段头key和字段体value组成字段头由一个变长32位整数表示字段体由具体的数据结构和数据类型决定。字段头格式(field_number 3) | wire_type-field_number字段序号-wire_type字段编码类型字段编码类型Type Meaning Used For0 Varint int32, int64, uint32, uint64, sint32, sint64, bool, enum1 64-bit fixed64, sfixed64, double2 Length-delimited string, bytes, embedded messages嵌套message, packed repeated fields3 Start group groups (废弃)4 End group groups (废弃)5 32-bit fixed32, sfixed32, float编码示例下面的编码以16进制表示示例1整数message Test1 {required int32 a 1;}a 150 时编码如下08 96 0108: 1 3 | 096 01:1001 0110 0000 0001- 001 0110 000 0001- 1001 0110- 150示例2字符串message Test2 {required string b 2;}b “testing” 时编码如下12 07 74 65 73 74 69 6e 6712: 2 3 | 207: 字符串长度74 65 73 74 69 6e 67- t e s t i n g示例3嵌套message Test3 {required Test1 c 3;}c.a 150 时编码如下1a 03 08 96 011a 3 3 | 203 嵌套结构长度08 96 01-Test1 { a 150 }示例4可选字段message Test4 {required int32 a 1;optional string b 2;}a 150, b不设值时编码如下08 96 01- { a 150 }a 150, b “aut” 时编码如下08 96 01 12 03 61 75 7408 96 01 - { a 150 }12 2 3 | 203 字符串长度61 75 74- a u t示例5重复字段message Test5 {required int32 a 1;repeated string b 2;}a 150, b {“aut”, “honey”} 时编码如下08 96 01 12 03 61 75 74 12 05 68 6f 6e 65 7908 96 01 - { a 150 }12 2 3 | 203 strlen(“aut”)61 75 74 - a u t12 2 3 | 205 strlen(“honey”)68 6f 6e 65 79 - h o n e ya 150, b “aut” 时编码如下08 96 01 12 03 61 75 7408 96 01 - { a 150 }12 2 3 | 203 strlen(“aut”)61 75 74 - a u t示例6字段顺序message Test6 {required int32 a 1;required string b 2;}a 150, b “aut” 时无论a和b谁的声明在前面编码都如下08 96 01 12 03 61 75 7408 96 01 - { a 150 }12 03 61 75 74 - { b “aut” }制代码还有什么编码风格花括号的使用(参考上面的proto文件)数据类型使用驼峰命名法AddressBook, PhoneType字段名小写并使用下划线连接person_info, email_addr枚举量使用大写并用下划线连接FIRST_VALUE, SECOND_VALU适用场景“Protocol Buffers are not designed to handle large messages.”。protobuf对于1M以下的message有很高的效率但是当message是大于1M的大块数据时protobuf的表现不是很好请合理使用。总结本文介绍了protobuf的基本使用方法和编码规则还有很多内容尚未涉及比如反射机制、扩展、Oneof、RPC等等更多内容需参考官方文档。标量类型列表proto类型 C类型 备注double doublefloat floatint32 int32 使用可变长编码编码负数时不够高效——如果字段可能含有负数请使用sint32int64 int64 使用可变长编码编码负数时不够高效——如果字段可能含有负数请使用sint64uint32 uint32 使用可变长编码uint64 uint64 使用可变长编码sint32 int32 使用可变长编码有符号的整型值编码时比通常的int32高效sint64 int64 使用可变长编码有符号的整型值编码时比通常的int64高效fixed32 uint32 总是4个字节如果数值总是比总是比228大的话这个类型会比uint32高效fixed64 uint64 总是8个字节如果数值总是比总是比256大的话这个类型会比uint64高效sfixed32 int32 总是4个字节sfixed64 int64 总是8个字节bool boolstring string 一个字符串必须是UTF-8编码或者7-bit ASCII编码的文本bytes string 可能包含任意顺序的字节数据protobuf 文件级别优化package IM.BaseDefine;option java_package “com.mogujie.tt.protobuf”;option optimize_for LITE_RUNTIME;。。。。option optimize_for LITE_RUNTIME;optimize_for是文件级别的选项Protocol Buffer定义三种优化级别SPEED/CODE_SIZE/LITE_RUNTIME。缺省情况下是SPEED。SPEED: 表示生成的代码运行效率高但是由此生成的代码编译后会占用更多的空间。CODE_SIZE: 和SPEED恰恰相反代码运行效率较低但是由此生成的代码编译后会占用更少的空间通常用于资源有限的平台如Mobile。LITE_RUNTIME: 生成的代码执行效率高同时生成代码编译后的所占用的空间也是非常少。这是以牺牲Protocol Buffer提供的反射功能为代价的。因此我们在C中链接Protocol Buffer库时仅需链接libprotobuf-lite而非libprotobuf。在Java中仅需包含protobuf-java-2.4.1-lite.jar而非protobuf-java-2.4.1.jar。SPEED和LITE_RUNTIME相比在于调试级别上例如 msg.SerializeToString(str) 在SPEED模式下会利用反射机制打印出详细字段和字段值但是LITE_RUNTIME则仅仅打印字段值组成的字符串;因此可以在程序调试阶段使用 SPEED模式而上线以后使用提升性能使用 LITE_RUNTIME 模式优化。在C环境下使用《 https://github.com/protobuf-c/protobuf-c 》《 https://github.com/nanopb/nanopb 》