Using "perfex" for P4 performance counters
THE ORIGINAL POST I WROTE LAST WEEKEND:
After spending much time playing perfex/lperfex, I have finally discovered the following (for Pentium 4 only!):
1) lperfex does NOT support P4 at all
2) it's very difficult to understand how to use perfex - its documentation is really poor
3) after looking at the Intel P4 manual and decoding the available perfex usage examples (none of which is about L2 cache miss), I have worked out two potential ways to get L2 cache miss counts:
#1:
There is an example in perfex.c to count L1 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000001 --p4pmv=0x1 some_program
Change it to the following to count L2 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000002 --p4pmv=0x1 some_program
On the P4 manual, this is a replay event (god knows what it is) and the replay metric is called 2ndL_cache_load_miss_retired, which has a footnote: 2nd-level misses retired does not count all 2nd-level misses. It only includes those references that are found to be misses by the fast detection logic and not those that are later found to be misses. [IA-32 Manual Vol. 3 A-35]
#2:
I finally got to know how to set the magic numbers (these numbers are not anywhere else on the internet):
perfex -e 0x0003f000/0x18020004@0x80000000 some_program
for RD_2ndL_MISS: read 2nd level cache miss (includes load and RFO).
perfex -e 0x0003f000/0x18000e04@0x80000000 some_program
for RD_2ndL_HITS + RD_2ndL_HITE + RD_2ndL_HITM: read 2nd level cache hit Shared/Exclusive/Modified (includes load and RFO).
I tested them, and it seems they are truely for the L2 cache. However, the number reports by #1 is always 3 ~ 5 times larger than that reported by #2. They are consistent with themselves as I change the test program's working set.
After spending much time playing perfex/lperfex, I have finally discovered the following (for Pentium 4 only!):
1) lperfex does NOT support P4 at all
2) it's very difficult to understand how to use perfex - its documentation is really poor
3) after looking at the Intel P4 manual and decoding the available perfex usage examples (none of which is about L2 cache miss), I have worked out two potential ways to get L2 cache miss counts:
#1:
There is an example in perfex.c to count L1 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000001 --p4pmv=0x1 some_program
Change it to the following to count L2 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000002 --p4pmv=0x1 some_program
On the P4 manual, this is a replay event (god knows what it is) and the replay metric is called 2ndL_cache_load_miss_retired, which has a footnote: 2nd-level misses retired does not count all 2nd-level misses. It only includes those references that are found to be misses by the fast detection logic and not those that are later found to be misses. [IA-32 Manual Vol. 3 A-35]
#2:
I finally got to know how to set the magic numbers (these numbers are not anywhere else on the internet):
perfex -e 0x0003f000/0x18020004@0x80000000 some_program
for RD_2ndL_MISS: read 2nd level cache miss (includes load and RFO).
perfex -e 0x0003f000/0x18000e04@0x80000000 some_program
for RD_2ndL_HITS + RD_2ndL_HITE + RD_2ndL_HITM: read 2nd level cache hit Shared/Exclusive/Modified (includes load and RFO).
I tested them, and it seems they are truely for the L2 cache. However, the number reports by #1 is always 3 ~ 5 times larger than that reported by #2. They are consistent with themselves as I change the test program's working set.
2 Comments:
You can set to monitor multiple events at the same time, as long as the counters don't conflict. Here are a few settings that I'm using:
1) L2 cache hits, L2 cache misses:
perfex -e 0x0003f000/0x18000e04@0x80000000 -e 0x0003f000/0x18020004@0x80000002 some_program
2) Retired instructions, data TLB misses, instruction TLB misses:
perfex -e 0x00039000/0x04000204@0x8000000C -e 0x00039000/0x02000204@0x80000000 -e 0x00039000/0x02000404@0x80000002 some_program
Reference:
1) IA-32 IntelĀ® Architecture Software Developer's Manual, Volume 3: System Programming Guide
ftp://download.intel.com/design/Pentium4/manuals/25366814.pdf
2) Some event specifiers for perfex on the Pentium
http://www.complang.tuwien.ac.at/anton/linux-perfex/pentium4
REPOSTING THE COMMENT (don't know why the original comment doesn't appear):
Basically you can monitor multiple events at the same time as long as these counters don't conflict. Here are a few settings I figured out. Note that they can be totally bogus - not verified with any authoritative.
1) L2 cache hits, L2 cache misses:
perfex -e 0x0003f000/0x18000e04@0x80000000 -e 0x0003f000/0x18020004@0x80000002
2) retired instructions, data TLB misses, instruction TLB misses:
perfex -e 0x00039000/0x04000204@0x8000000C -e 0x00039000/0x02000204@0x80000000 -e 0x00039000/0x02000404@0x80000002
Reference:
1) IA-32 IntelĀ® Architecture Software Developer's Manual, Volume 3: System Programming Guide
ftp://download.intel.com/design/Pentium4/manuals/25366814.pdf
2) Some event specifiers for perfex on the Pentium 4
http://www.complang.tuwien.ac.at/anton/linux-perfex/pentium4
Have fun and correct me if anything is wrong.
Post a Comment
<< Home