When dealing with phase, we first must be aware of some facts regarding A) the physical nature of sound and B) the way microphones work. I was not totally aware of these facts until I found myself dealing with phase issues and decided to understand them. Admittedly, this was not an easy task. I found these facts to be somehow hidden within the dedicated literature, inaccurately described or else misrepresented in different ways, and sometimes completely ignored. What follows is the result of a personal research both theoretical and practical in the use of two or more microphones.
The first fact to be aware of you might already know: phase is frequency-dependant.
There is no such thing as a "general phase problem" pertaining to a combination of signals generated by multiple microphones. For a given combination, a phase problem always occurs at a certain fundamental frequency and at higher, mathematically related, frequencies (depending on the type of microphones used). As previously said, sound waves have peaks and dips of pressure: physical space (distance) between them measures differently for each frequency because each frequency has its own wavelength. In the example above, when a peak at the second microphone corresponds to a dip at the first microphone, the result is a total cancellation of one fundamental frequency (plus higher frequencies). But in any case like this, there will be a different wavelength (i.e. frequency) for which both microphones simultaneously pick up a peak (or a dip), and the result is an increase in volume for that frequency. In other words, any time there is cancellation at one fundamental frequency, there will be a boost at another fundamental frequency.
This is very important to be aware of: phase is always a risk and — at the same time — an opportunity to carve a desired tonal shape using your main recording tools: the microphones. Frequencies are canceled while frequencies are enhanced.
There is a second fact we must be aware of: microphones don't translate sound pressure in the same way. No, I am not talking about different frequency responses. Maybe you know this as well, but are you sure you know the whole story?
The types of microphones widely used nowadays are condensers, dynamics, and ribbons.
In a condenser microphone, the transducing element, called a capsule, consists of a supporting ring to which are attached an extremely thin, circular sheet of metal-sputtered plastic material under tension, called a diaphragm, and a thicker, fixed, perforated metallic plate behind that (if you follow the direction of sound wave). (In the majority of condenser capsules in use today, two diaphragms are present, one on each side of the backplate. For the purpose of this explanation we can just consider the front one.)
The two elements, diaphragm and backplate, are separated by an air gap and are polarized by an electrical charge provided, effectively constituting the two arms of a capacitor (condenser is an old term for capacitor). The thin diaphragm involved in the periodic variation of air pressure of the sound wave vibrates together with it. With positive pressure, the diaphragm is pushed against the backplate, the distance between them shortens and so the capacitance of the system is varied. The other way around, the diaphragm is pulled by de-compression; the distance from the backplate increases and the capacitance of the system varies in the opposite verse. Said oscillation of the capacitance is electrically used to generate the variation of voltage at the output of the microphone. Note that the amplitude of the output signal is highest when the diaphragm is at its closest or farthest position from the backplate. Therefore in a condenser microphone the level of the signal is directly proportional to the air pressure: the more pressure (or de-pressure) the more signal.
With dynamic and ribbon mics, transduction of air pressure into electric voltage is accomplished using electromagnetic induction. Similarly to a condenser, also in dynamic microphones a diaphragm is present. The whole system though is heavier — i.e. slower and "less accurate" — because attached to the diaphragm is a coil of thin copper wire. Said coil of wire surrounds a fixed, strong magnet. According to the principles of electromagnetic induction, voltage is induced on the coil when it moves (remember: it's attached to the diaphragm) and interferes with the magnetic field emanated by the magnet.
Note that movement is the necessary condition for voltage to be induced on the coil: if no movement, then no output signal. Here the amplitude of the output signal is highest when the diaphragm is moving at its fastest speed. What you must understand is that the fastest speed of the diaphragm takes place exactly in between two opposite peaks of pressure, i.e. when there's no pressure on the diaphragm. This is exactly when the output of a condenser microphone is at zero. Therefore, in a dynamic microphone the level of the signal is inversely proportional to the air pressure: the more pressure (or de-pressure) the less signal.
Ribbon microphones operate in a similar fashion, also according to the principles of electromagnetism. A conductive metallic element as a very thin ribbon is kept under tension, free to oscillate in between the two poles of a magnet and connected to the output pins. Voltage is present at the output pins when the ribbon, hit by air pressure, moves within the magnetic field and in doing so gets voltage induced onto itself. Exactly like a dynamic capsule (the principle is the same), the movement of the ribbon is the necessary condition for voltage to be present at the output. And also in this case the output level is inversely proportional to the pressure on the diaphragm.
Okay, let's categorize dynamic and ribbon microphones as electromagnetic microphones.
Following what's been said above, let's plot the output of a condenser microphone and the output of an electromagnetic microphone as the same sound pressure varies on their diaphragms: we'll obtain two resulting waveforms which are 90° out of phase.
You can verify this by yourself: set up two microphones, one condenser and one dynamic (or ribbon) very close to each other, as coincident as possible, at a few inches from a P.A. speaker reproducing a single frequency, say a 100 Hz tone. Record on two different tracks the output of the mics (use preamps of the same brand/type). Assuming you are recording into a DAW, zoom in a lot and look very close at the two waveforms: you'll see that the electromagnetic precedes the condenser (look at the peaks): it picked up the same air pressure but translated that "earlier." The two output signals ARE NOT IN PHASE even though the two microphones are perfectly aligned in front of the speaker!
The plot provides very useful information about what to do next. Never forget that phase is frequency dependant! The plot tells you where to position microphones in order to tune them (i.e. obtain perfect phase) at a specific frequency you want to enhance or else get rid of. It tells you that if they transduce the same way (both electromagnetics or both condensers), you could position them as coincident as possible. It also tells you that if they transduce the different way, you should firmly avoid the "as coincident as possible" positioning. If for any given frequency, electromagnetic transduction "precedes" the capacitive one, the next thing to know is that you must place the condenser closer to the source, ahead of the electromagnetic in order to compensate for the phase shift. At what distance from the electromagnetic? How much should the two microphones be spaced apart? The plot says one quarter wavelength of the target frequency.