Performance counter patch for Linux
Some of you have already installed this, but if not, here's the link to the perfctr patch. If you have any other links to performance-counter related stuff, please add it in comments to this post.
A blog for the PLASMA group, and friends.
2 Comments:
Here are some more performance counter related tools. Not sure which ones are good... Brink and Abyss, PerfSuite, and OProfile.
After spending much time playing perfex/lperfex, I have finally discovered the following (for Pentium 4 only!):
1) lperfex does NOT support P4 at all
2) it's very difficult to understand how to use perfex - its documentation is really poor
3) after looking at the Intel P4 manual and decoding the available perfex usage examples (none of which is about L2 cache miss), I have worked out two potential ways to get L2 cache miss counts:
#1:
There is an example in perfex.c to count L1 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000001 --p4pmv=0x1 some_program
Change it to the following to count L2 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000002 --p4pmv=0x1 some_program
On the P4 manual, this is a replay event (god knows what it is) and the replay metric is called 2ndL_cache_load_miss_retired, which has a footnote: 2nd-level misses retired does not count all 2nd-level misses. It only includes those references that are found to be misses by the fast detection logic and not those that are later found to be misses. [IA-32 Manual Vol. 3 A-35]
#2:
I finally got to know how to set the magic numbers (these numbers are not anywhere else on the internet):
perfex -e 0x0003f000/0x18020004@0x80000000 some_program
for RD_2ndL_MISS: read 2nd level cache miss (includes load and RFO).
perfex -e 0x0003f000/0x18000e04@0x80000000 some_program
for RD_2ndL_HITS + RD_2ndL_HITE + RD_2ndL_HITM: read 2nd level cache hit Shared/Exclusive/Modified (includes load and RFO).
I tested them, and it seems they are truely for the L2 cache. However, the number reports by #1 is always 3 ~ 5 times larger than that reported by #2. They are consistent with themselves as I change the test program's working set.
Post a Comment
<< Home